Building enterprise AI agents @Microsoft. Author of “Behind the Interface”. Built gettenai.com, iheartbooks.club

Agentic AI products are going to switch to Grok 4.5. It’s at least 2x faster and an order of magnitude cheaper (cheaper pricing + efficient trajectories) at the same quality as flagship models from A/ and OAI. If your company isn’t making the move, you are going to lose @SpaceXAI
1
2
23
Why @AnthropicAI uses @fin_ai for their own website’s chatbot? And chatbot doesn’t even use Anthropic models? Fin’s website says they use their own proprietary models. I guess SaaS not dead yet
1
24
OpenAI just copied my project 😭 Sceptical about sharing your data with @OpenAI? Use PFA instead: github.com/gauravtendolkar/P… 100% local LLM - Data stays on your device #LocalAI #PersonalFinance
1/6 I just asked my bank accounts: “If I cut restaurant expenses to once a month and use that for a travel fund, how many months until I can afford a 3 day Disneyland trip?” It did the math. Across every account! You can do this too - 100% privately. 🧵
1
32
Can you guess which one is @Alibaba_Qwen Qwen 3.5 2B and which one is @GoogleDeepMind Gemma 4 2B? Running locally on @get_ten_ai iOS app
2
1
131
1/6 I just asked my bank accounts: “If I cut restaurant expenses to once a month and use that for a travel fund, how many months until I can afford a 3 day Disneyland trip?” It did the math. Across every account! You can do this too - 100% privately. 🧵
5
125
6/6 It’s fully open-source and free. GitHub 👇 github.com/gauravtendolkar/P… ⭐ Star it if you’d use this 🔁 Repost if you know someone paying for money management apps 💬 Would love contributions
1
1
39
5/6 That last part is the whole point. YNAB, Copilot, Monarch - your transactions run through their servers. You’re trusting a privacy policy. With PFA, the AI is local. Your bank data never leaves your computer. Not because of a promise. Because it physically can’t.
21
4/6 Setup takes minutes and you stay in full control: → Connect any bank via SimpleFin (read-only access) → A local LLM (Qwen 3.5) runs entirely on YOUR machine → Ask anything in plain English → Get real answers across all your accounts No subscription. No one sees your data.
44
3/6 So I built PFA - an open-source personal finance agent. Not a chatbot. Not a dashboard. An AI that actually reasons over what you’ve spent, where, and when - and answers the question YOU have. Free and Private.
21
2/6 No finance app can do this. Not YNAB. Not Copilot. Not Monarch. Not Mint. They’re all just dashboards with prettier charts. Pre-programmed views for questions nobody actually asks. Yours is always something specific. Something personal. Something they never built a tab for.
34
Money apps @ynab @copilotmoney @mint reading your transactions & charging you for it! Sharing PFA - open-source finance agent that runs on-device w/ @Alibaba_Qwen 3.5. Private & Secure. Connect any bank (read-only via SimpleFin). Ask any questions Follow - dropping on GitHub soon
1
44
Even in 2026, even with today’s coding agents, building a web based text editor with support for pages (simple pages like word or google docs!!!) is insanely complicated. The core problem being ability to measure text rendering precisely. Thank you @_chenglou 🙏🏻🫡
My dear front-end developers (and anyone who’s interested in the future of interfaces): I have crawled through depths of hell to bring you, for the foreseeable years, one of the more important foundational pieces of UI engineering (if not in implementation then certainly at least in concept): Fast, accurate and comprehensive userland text measurement algorithm in pure TypeScript, usable for laying out entire web pages without CSS, bypassing DOM measurements and reflow
38
I feel like I have unlocked new powers! Forever running research agent running while I am sleeping. Just one hour, just 4K tokens and finding big gains!
Three days ago I left autoresearch tuning nanochat for ~2 days on depth=12 model. It found ~20 changes that improved the validation loss. I tested these changes yesterday and all of them were additive and transferred to larger (depth=24) models. Stacking up all of these changes, today I measured that the leaderboard's "Time to GPT-2" drops from 2.02 hours to 1.80 hours (~11% improvement), this will be the new leaderboard entry. So yes, these are real improvements and they make an actual difference. I am mildly surprised that my very first naive attempt already worked this well on top of what I thought was already a fairly manually well-tuned project. This is a first for me because I am very used to doing the iterative optimization of neural network training manually. You come up with ideas, you implement them, you check if they work (better validation loss), you come up with new ideas based on that, you read some papers for inspiration, etc etc. This is the bread and butter of what I do daily for 2 decades. Seeing the agent do this entire workflow end-to-end and all by itself as it worked through approx. 700 changes autonomously is wild. It really looked at the sequence of results of experiments and used that to plan the next ones. It's not novel, ground-breaking "research" (yet), but all the adjustments are "real", I didn't find them manually previously, and they stack up and actually improved nanochat. Among the bigger things e.g.: - It noticed an oversight that my parameterless QKnorm didn't have a scaler multiplier attached, so my attention was too diffuse. The agent found multipliers to sharpen it, pointing to future work. - It found that the Value Embeddings really like regularization and I wasn't applying any (oops). - It found that my banded attention was too conservative (i forgot to tune it). - It found that AdamW betas were all messed up. - It tuned the weight decay schedule. - It tuned the network initialization. This is on top of all the tuning I've already done over a good amount of time. The exact commit is here, from this "round 1" of autoresearch. I am going to kick off "round 2", and in parallel I am looking at how multiple agents can collaborate to unlock parallelism. github.com/karpathy/nanochat… All LLM frontier labs will do this. It's the final boss battle. It's a lot more complex at scale of course - you don't just have a single train. py file to tune. But doing it is "just engineering" and it's going to work. You spin up a swarm of agents, you have them collaborate to tune smaller models, you promote the most promising ideas to increasingly larger scales, and humans (optionally) contribute on the edges. And more generally, *any* metric you care about that is reasonably efficient to evaluate (or that has more efficient proxy metrics such as training a smaller network) can be autoresearched by an agent swarm. It's worth thinking about whether your problem falls into this bucket too.
1
83
Gaurav T retweeted
Most on-device AI apps aren’t useful. MLX by design hogs memory and iOS crashes the app on longer conversations. Ten AI is the most efficient private on-device AI super app built on a custom fork of llama.cpp. Qwen 3.5 4B with vision will only consume upto 1 GB memory with Ten AI
11
1
11
151
The @Apple’s app review process taking longer than building the app 😓
34
Gaurav T retweeted
🥝 Meet Kimi K2.5, Open-Source Visual Agentic Intelligence. 🔹 Global SOTA on Agentic Benchmarks: HLE full set (50.2%), BrowseComp (74.9%) 🔹 Open-source SOTA on Vision and Coding: MMMU Pro (78.5%), VideoMMMU (86.6%), SWE-bench Verified (76.8%) 🔹 Code with Taste: turn chats, images & videos into aesthetic websites with expressive motion. 🔹 Agent Swarm (Beta): self-directed agents working in parallel, at scale. Up to 100 sub-agents, 1,500 tool calls, 4.5× faster compared with single-agent setup. - 🥝 K2.5 is now live on kimi.com in chat mode and agent mode. 🥝 K2.5 Agent Swarm in beta for high-tier users. 🥝 For production-grade coding, you can pair K2.5 with Kimi Code: kimi.com/code - 🔗 API: platform.moonshot.ai 🔗 Tech blog: kimi.com/blogs/kimi-k2-5.htm… 🔗 Weights & code: huggingface.co/moonshotai/Ki…
762
1,941
15,608
7,362,433
Time to cook👨‍💻 This new Qwen TTS model looks 🔥
Qwen3-TTS is officially live. We’ve open-sourced the full family—VoiceDesign, CustomVoice, and Base—bringing high quality to the open community. - 5 models (0.6B & 1.8B) - Free-form voice design & cloning - Support for 10 languages - SOTA 12Hz tokenizer for high compression - Full fine-tuning support - SOTA performance We believe this is arguably the most disruptive release in open-source TTS yet. Go ahead, break it and build something cool. 🚀 Everything is out now—weights, code, and paper. Enjoy. 🧵 Github: github.com/QwenLM/Qwen3-TTS Hugging Face: huggingface.co/collections/Q… ModelScope: modelscope.cn/collections/Qw… Blog: qwen.ai/blog?id=qwen3tts-011… Paper: github.com/QwenLM/Qwen3-TTS/… Hugging Face Demo: huggingface.co/spaces/Qwen/Q… ModelScope Demo: modelscope.cn/studios/Qwen/Q… API: alibabacloud.com/help/en/mod…
45
We seem to be at GPT 2 era of 3D world generation!
The World API is live. Generate persistent, explorable 3D worlds from text, images, and video. Integrate them directly into your products.
47
If you are in college & looking to get into robotics (as an entrepreneur or otherwise), bookmark this right now! This is the most comprehensive post on the current landscape of robotics.
Spend an hour reading this weekend and I think you’ll know more about robotics than 99% of people, including some people who invest in robotics. notboring.co/p/robot-steps
40