Live now: our Vision & OCR Track from AI Engineer World's Fair 2026.
A model that counts 32 white squares on part of a chessboard. A file format that stores a table as a pile of line segments. Ten turkeys on the roof of a Tesla.
Thesis: the models can see. They are still learning to look.
piped.video/watch?v=RQi7x-na…
- Building the Document Context Layer for AI Agents:
@jerryjliu0, LlamaIndex
- Skill issue: stop deploying vision language models, use them with Skills:
@mervenoyann, Hugging Face
- Modality Misalignment and Originality Attribution in Short-Form Video: Aditya Gautam, Meta
- From Ingestion to Agents: How AI Teams Build on Document Intelligence: Adit Abraham, Reducto
- The Best Models Still Reason Like Toddlers:
@andrewdai, Elorian
- You're Not Thinking Big Enough: Rebuilding Food Systems with AI Agents:
@cbmenefee, Firecrawl
- From VLM/VLA's to Embodied Agents:
@ArmenAgha, Perceptron AI
- From Scratch to SOTA: Training a 3B State-Space Vision Model:
@fewshotlearner, Sarvam