I’m building something awesome. 👀
The inspiration came while on holiday with my kids—but this is the next chapter of my AI-native device project, Vault.
I’ve tested the stack with current SOTA open-weight LLMs, and it works. The ultra-compact enclosure for my custom Jetson Orin Nano Super build is now printing. Fully parametric, too—no Astra operating Blender involved. 😉
Probably my most ambitious after-hours Physical AI project yet. Still need to add a physical kill switch… just in case! 😅
I’m fascinated by how far I can push this.
Where it started:
callstack.com/blog/building-…#PhysicalAI#OpenSource#Hardware
Any agent can test iPhone Duo apps for you now. No need to do it by hand.
Folded, half-open, open, all without stealing your focus or other agents sessions.
Just give them @agent_device
How does quantization affect APEX performance? 🔬
Artur Morys-Magiera @artus9033 and I shared our research at @AgentConf . We hope our findings help the community make more informed decisions about quantization.
🎥 piped.video/XLoY8PpYXNU
344 runs. 25 days. Quantization is humbling 😅
Apex-31B BF16: 57% → 86.5% MMLU after fixing the chat template + eval harness. Same weights. Same 400 questions. 🤯
Then the payoff: IQ4_XS gave us 3.4× decode throughput and a 73% smaller file vs BF16 on DGX Spark. 🚀
@artus9033 and I talked about all of this at @AgentConf , organized by @callstackio.
For me, it’s already one of the most exciting AI conferences in Poland! 🇵🇱🔥
I’m still so surprised that, besides being a speaker at @code_europe, I WON a @PlayStation 5! 🎮🔥
Huge thanks to @softswiss for the generous prize in their lottery!
The funniest part? When I first got the “Congratulations, you won!” email, I was convinced it was phishing. 😅
Turns out it was very real. What a great surprise after Code Europe!
Last preparations yestarday. Today @AgentConf begins! I will be at APEX and Agent Device booth at 12. We can talk about our latest achievements in AI and agentic workflows. Meet me there!
Releasing free unlimited React Native cloud Simulators
An OSS CLI that hosts simulators on Github and lets you or your agent run @agent_device to test things out. Available for private repos too.
Try out the alpha today at reactnativefeel.com/sim or install right away w/ `npm i -g native-sim`
GPT-6 Astra vs Claude Fable 5.1 on RN Evals.
134 tasks. 10 runs each. 2,680 fully judged evaluations.
Published overall scores:
Astra: 86.4%
Fable: 89.0%
Categories: Navigation, Animation, Async State, Lists & React Native APIs.
A narrow lead for Fable. One run doesn’t tell the whole story.
rn-evals.vercel.app/
Agents Commander 0.1.5 is out.
Now MIT-licensed: up to 100 agent panels, OpenCode, local agent-to-agent messaging, and opt-in capture with reviewed training-data export.
Try the offline demo:
npx agents-commander@0.1.5 --demo
npmjs.com/package/agents-com…
React Native Evals are out with a bunch of new models:
> Claude Opus 5 only +4 pp over Opus 4.6 at ~2× the cost
> Grok 4.6 scores 12 pp higher than Grok 4 at ~half the cost
> Ox Alpha hits 81% at $0, basically free Opus 4.7 territory