RL envs @vibrantlabsai evals @ragas_io 💻 Building devtooling and maintain oss⚡️

🇮🇳
If there is a possibility, it is a reality!
3
4
56
This past weekend @fhackdroid and me got an opportunity to talk at IICT 2026 @compiler_tech hosted at @iiscbangalore on waspy, a python to wasm compiler, our progress so far, things that work now, and things being worked upon... 🎉
1
1
15
411
We also introduced waspy playground...
We've a new Python to WASM playground on waspy website now, where you can try diff python programs and see it's WAT, as well as functions to interact with. First presented yesterday at CompilerTech conference at IISC, Bengaluru with @fhackdroid Come check this...
1
1
48
...and we did not realise, when the limits were risen...
99
We've a new Python to WASM playground on waspy website now, where you can try diff python programs and see it's WAT, as well as functions to interact with. First presented yesterday at CompilerTech conference at IISC, Bengaluru with @fhackdroid Come check this...
1
5
300
This uses current state of waspy to compile python to wasm, so not all cpython modules are currently supported... You can see all the modules supported currently here: anistark.github.io/waspy/mod… ...and we keep this updated as we work more on it... This is based on py 3.12 so perhaps we might be missing some modules, but so far, this is the plan till v1 🔥
1
1
42
Play around with the playground here: anistark.github.io/waspy/pla… Let me know if you run into issues... or better yet, help us know by filing it on : github.com/anistark/waspy
2
27
Kumar Anirudha retweeted
RL environments are fundamentally a form of data. Like pretraining data and SFT trajectories, they can push the frontier while they’re unsaturated. Once they saturate, adding more of the same stops helping
Eventually RL envs will not be needed
2
1
33
1,664
Kumar Anirudha retweeted
We'll be spending a lot more time trying to understand the outputs of language models. A few thoughts, tips & tricks: Writing. Something I've had success with: Ask your LLM to explain something in ASD-STE100, it's a controlled language specification originally developed for aerospace maintenance documentation. LLMs well-versed in this language and it comes with heavy constraints on clean writing style that I often find a lot more readable. Sometimes I've tried to soften it a bit e.g. ask for "80% of the way to ASD-STE100" because the spec is quite stringent. But even better: Diagrams / images. Instead of writing, ask your LLM to create a diagram. These can be a lot easier to process, parse, and understand. But even better: Web pages. Ask for output "in HTML" to get a beautiful, interactive webpage. LLMs are getting really good at frontend and can create beautiful experiences, animations, etc. But even better: Explainer videos. The output format I am most bullish on is fully custom / bespoke explainer videos generated on any arbitrary topic. Experiment with things like "Create a 3b1b style video explainer on X. Use my ElevenLabs API key for audio narration". (you'd need an API key for the latter or you can ask your LLM to find you decent free alternatives that use your local compute). This is actually starting to work! In summary: - As LLMs get better, they will do more and more of the legwork autonomously, and a lot more of our work will rise up the abstractions into oversight and understanding. - Luckily, LLMs can help here too because as intelligence and code are increasingly abundant, you can ask for large, custom, discardable software artifacts (e.g. web apps, video explainers) that would have never made sense to create before. Push the boundaries here and you'll be surprised.
1,559
6,368
54,213
7,520,203
Suddenly, we're back to clicking on every vercel link that pops up ?
112
When LLM doesn't understand what "realistic" means...

ALT I Get It Bro GIF by Dead Meat James

83
Kumar Anirudha retweeted
This is something we have been exploring at @VibrantLabsAI and as models get better their ability to do this will keep getting better but that brings me to the question, is there any benchmark that already tracks this?
⏱️⚙️Introducing *AutoBenchmark* 📊🏁 - Creating benchmarks automatically - Benchmarking benchmark creation - Studying the role & impact of humans in the loop Blog post: facebookresearch.github.io/R… Key takeaways: 1) We find human-agent collaboration gives big wins over agents alone - fine-grained feedback in ideation stage crucial - autobench can be used to measure this in the future with stronger agents 2) We show that it's possible to make *autoresearch benchmarks* for AI research using this recipe – full recursive improvement loop! 3) Important Ingredients: Autobenchmark creation works best with feedback from two sources: benchmark solvers + external verifiers (human+AI). 🧵1/5
1
1
1
80
Curious to know if there's any analytics data around @AnthropicAI Fable usage since 5.5 class models... @trq212 would you know?
97
The most competitive age of tech...
2
80
Haiku still looking for invite to the 5.5 party 🥲
3
141
"Personalisation" in every products is gonna take a giant leap forward.

ALT Spider-Man No GIF by nounish ⌐◨-◨

1
99
Sonnet 5.5 better than Opus 5.5 on agentic coding 😮
Replying to @claudeai
Sonnet 5.5 improves on Sonnet 5 across benchmarks, in some cases dramatically. It’s a faster, lower-cost complement to Claude Opus 5.5, strongest at well-scoped everyday tasks, fixing bugs, and creating polished documents, slides, and spreadsheets.
1
2
153
Had a great time talking about building 3D worlds… So great to find folks working on incredible stuff in the area… at @IndiaFOSS 2026 ⚡️
1
2
23
311
Find out more about runek.nullorder.org here.
2
38