Scaling Applied Intelligence • Animation Bench • hello@physera.ai

Pinned Tweet
Introducing Animation Bench, our first benchmark to evaluate frontier models on web animation reconstruction. Coding agents can already recreate visually plausible web animations, this work evaluates if they can reconstruct behavior of a web page.
4
14
83
17,981
Physera retweeted
the most interesting fact about animation bench release is i had calls with six researchers from diff labs / neolabs in last two days and there first words were "its just crazy that no one has thought about this capability gap earlier". we are coming up with more announcements soon!
3
3
82
3,320
We evaluated GPT-6.1 Sol on Animation Bench and it is the best model to perform overall as well as on motion consistency. It is substantially better at recovering timing, interaction state and coordinated movement while retaining low cost for Sol.
Introducing Animation Bench, our first benchmark to evaluate frontier models on web animation reconstruction. Coding agents can already recreate visually plausible web animations, this work evaluates if they can reconstruct behavior of a web page.
2
3
10
3,573
Today we introduce Animation Bench, a benchmark for how well frontier multimodal coding agents can recreate real web animations. We gave frontier models 48 animations from 32 production websites, covering things like scroll, hover, drag, click, cursor movement and state changes. We show that while frontier models are getting pretty good at recreating what a website looks like (the layout), there is still room for improvement in recreating how it moves. Across 192 reconstructions, visual similarity was better than motion consistency in 177 of them. The median reconstruction captured only 63% of the reference motion. 38% captured less than half.
Introducing Animation Bench, our first benchmark to evaluate frontier models on web animation reconstruction. Coding agents can already recreate visually plausible web animations, this work evaluates if they can reconstruct behavior of a web page.
2
1
12
1,403
We are happy to announce our new benchmark, and I am proud of the work done by our team. As we move towards RSI, the benchmarks used to measure how well AI models are performing must reflect significant advancements and continue to raise the bar. We will be releasing many ahead.
Introducing Animation Bench, our first benchmark to evaluate frontier models on web animation reconstruction. Coding agents can already recreate visually plausible web animations, this work evaluates if they can reconstruct behavior of a web page.
1
3
20
1,587
Physera retweeted
Today, I am excited to announce Animation Bench. Our first benchmark to evaluate frontier models on web animation reconstruction. Frontend generation is a common commercial use case for coding agents. In addition to static layout it contains transitions and interactions triggered by dynamic behaviours. We evaluated four frontier models on 48 tasks from real websites across three axes: 1. Visual Similarity 2. Motion Consistency 3. Layout Correctness. We find that motion consistency is often incomplete or missing.
Introducing Animation Bench, our first benchmark to evaluate frontier models on web animation reconstruction. Coding agents can already recreate visually plausible web animations, this work evaluates if they can reconstruct behavior of a web page.
22
16
256
20,529
Introducing Animation Bench, our first benchmark to evaluate frontier models on web animation reconstruction. Coding agents can already recreate visually plausible web animations, this work evaluates if they can reconstruct behavior of a web page.
4
14
83
17,981
Across all 4 models, mean visual similarity exceeds mean motion consistency. At the task level, visual exceeds motion in 177 of 192 reconstructions.
1
6
276
Physera retweeted
we evaluated @grok-4.7 on an internal sample of 22 knowledge-work tasks across engg, bizops, finance and healthcare. grok-4.7 is at #3 on the table, costed <$5. - really strong in multi-document policy reasoning and precedence handling. - good at large structured audits across csv, json, markdown, and spreadsheets. - consistently produced the requested deliverable files.
Grok 4.7 is here. It's a notable improvement over Grok 4.6 at the same price and speed.
6
8
96
20,893
Physera retweeted
Deepseek V4.1 Flash is nearly SOTA on KnowledgeWork eval (15/21 passes) being 98% cheaper than Fable 5.1 (18/22 passes). The most surprising result was GPT-6 Astra (11/22 passes). We evaluated 11 frontier models on same internal sample of 22 tasks across Engg, BizOps, Finance and Healthcare. 🧵Analysis from agent trajectories
🚀 Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient. 🔹 Introducing the smallest model in our new architecture family, with native visual understanding. 🔹 Designed for greater capability, faster inference, higher throughput, and scaling to larger models. 1/6
2
6
71
11,088
Physera retweeted
We just evaluated Gemini-3.8-Flash on our @PhyseraAI - TB bench. Our analysis on 5 random tasks from the set: 1. It is excellent at deriving and implementing a coherent numerical method. It turned an internally consistent mathematical model into working code, especially when it can derive a checkable optimum. 2. It is inconsistent on edge-case semantics. It knows the individual primitives, but assemble them in the wrong order when a specification has interacting edge cases. 3. It tends to overbuild static-analysis solutions while missing the hardest coverage cases. 4. In multiple tasks long trajectories do not imply better outcomes. More exploration became speculative scope expansion rather than targeted verification. I think it is promising low-cost choice for numerical/scientific coding / transformations with crisp formulas / tasks where it can independently check residuals or invariants but had fallbacks for production shell tooling / static analysis / clinical derivations / compliance-style work.
Introducing Gemini 3.8, our best reasoning & coding model yet. By leveraging long-running agentic loops, we’re building on the momentum of 3.7 Flash from just three weeks ago to release two new 3.8 variants: Meet Gemini 3.8 Flash and Gemini 3.8 Flash Cyber 🧵
5
5
69
11,637
Physera retweeted
We have evaluated 0xAlpha (GLM-5.3-Flash) on our @PhyseraAI - TB bench, an internal TerminalBench3.0-style benchmark. Our analysis: - It is strong on shell/devops reliability work. It repeatedly produced a correct fail-closed reporting script / handling malformed manifests / relative paths / atomic report updates and temporary-directory cleanup. - It is strong on spec-driven implementation. It repeatedly solved SVG stop ordering / duplicate offsets, repeat/reflect behavior / sRGB vs linearRGB interpolation and opacity. - In multiple tasks, it got timed out three times at its 2-hour limit. It made substantial progress, but its partial implementations still missed dynamic-import cases such as mapping/star-import handling. Overall, it's a good model sir. We have got enough evidences passing the same tasks where frontier models have scored low. Bullish!
Introducing GLM-5.3-Flash - Leading capabilities at a highly competitive price - Natively multimodal with a 1M-token context window - A 320B-A18B model released under the MIT License - Previously previewed as Ox Alpha, running entirely on Chinese AI chips Blog: z.ai/blog/glm-5.3-flash Available now across all official platforms: Weights: huggingface.co/zai-org/GLM-5… API: docs.z.ai/guides/llm/glm-5.3… Coding Plan: z.ai/subscribe ZCode: zcode.z.ai/en Chat: chat.z.ai AutoClaw: autoclaw.z.ai
3
6
73
9,486
Physera retweeted
Today we introduce Physera to the world! Physera is an applied research and product lab working at the intersection of model efficiency and behavioural simulations. We are rethinking each layer of AI stack from first principles. 1. We believe there has been no better time to scale capabilities of frontier models with efficient architectures. 2. Simulating human decision-making with high-fidelity and multimodal environments. 3. Translating human judgment into models. We think the frontier of prediction has always been gated by the number of controlled variables we can simulate. We're looking for thoughtful folks to help shape this vision. Happy to chat!
Today we introduce Physera, a research and product lab rethinking applied intelligence. We are working at the intersection of model efficiency and behavioural simulations while building environments that are multimodal. We are a team of applied researchers and engineers who believe that the important problems in AI today are not about capability but making that capability reliably useful across multimodality. We are heads down building systems that perceive, reason, and decide as humans do, under the constraints humans face. To learn more or collaborate: physera.ai/?v
85
28
419
53,948
Physera retweeted
Excited to introduce @PhyseraAI. Physera is a research and product lab rethinking applied intelligence. We’re building across three fronts: 1. Making AI systems efficient under real-world cost and latency constraints 2. Simulating human decision-making with high-fidelity, multimodal environments 3. Turning model understanding into measurable, reliable outcomes AI progress isn’t just about capability, it’s about making that capability useful, predictable, and deployable. We’re rethinking the stack from first principles. We’re assembling a small group of thoughtful builders around a mission to create something meaningfully different. If that sounds interesting, I’d love to chat.
Today we introduce Physera, a research and product lab rethinking applied intelligence. We are working at the intersection of model efficiency and behavioural simulations while building environments that are multimodal. We are a team of applied researchers and engineers who believe that the important problems in AI today are not about capability but making that capability reliably useful across multimodality. We are heads down building systems that perceive, reason, and decide as humans do, under the constraints humans face. To learn more or collaborate: physera.ai/?v
9
4
145
39,038
Introducing @PhyseraAI, an applied research and product lab working on the next layer of useful intelligence. We are focused on efficient AI systems, high-fidelity behavioral simulations, and turning model understanding into reliable outcomes.
Today we introduce Physera, a research and product lab rethinking applied intelligence. We are working at the intersection of model efficiency and behavioural simulations while building environments that are multimodal. We are a team of applied researchers and engineers who believe that the important problems in AI today are not about capability but making that capability reliably useful across multimodality. We are heads down building systems that perceive, reason, and decide as humans do, under the constraints humans face. To learn more or collaborate: physera.ai/?v
1
4
18
4,502
Today we introduce Physera, a research and product lab rethinking applied intelligence. We are working at the intersection of model efficiency and behavioural simulations while building environments that are multimodal. We are a team of applied researchers and engineers who believe that the important problems in AI today are not about capability but making that capability reliably useful across multimodality. We are heads down building systems that perceive, reason, and decide as humans do, under the constraints humans face. To learn more or collaborate: physera.ai/?v
7
10
145
110,038