MTS | Agentic Engineering | OpenSkills Creator | OSS

Glasgow, Scotland
LLM Progression 2026 Opus 5.5 in Cowork from Mobile Raw Prompt 👇
1
10
612
This is so much fun, need to get back to playing with the technology and not only building Voice prompted this: “Can you make me a video of LLM progression, like models that is from January of this year up until now, and what the trajectory is looking like, the capabilities and models as well? Like the capabilities of models too? I want to make it really like bang, bang, boom, whoa, here we go, kind of amazing innovation. Are you ready for this kind of thing? I think you'd use vMotion, 3GS, a bunch of stuff you've got. Use whatever you want.”
1
158
What the actual f*** This is too damn good! Everything will be generated The amount of cross cutting intelligence needing to make this real God what a wonder
Replying to @pleometric
that’s pretty cool! i also gave it a try with a different style direction and the result (after a few rounds of tweaking) is impressive
2
2
94
20,196
These Opus 5.5 videos are absolute gold
one-shotted this music video using opus 5.5 (life of coderabbit's rabbit)
1
589
I love this future of gaming
Introducing Agora-2, our next-generation multi-agent world model. Agora-2 supports up to 20 humans and agents interacting inside a shared environment, all simulated in real time. Our multiplayer research preview is available to try right now!
1
4
1,343
Three.js is going to become the next biggest game engine
Opus 5.5 has once again destroyed my expectations for games and 3D modeling. A full survival game like The Forest, in three.js, in 1 day of continual work. A 2 km island carved by simulated erosion. 44,000 trees, every one generated in code and wrapped in real photoscanned bark. Six real Pacific Northwest species. The sun and moon sit where they really would at 49 degrees north in late September. Thunder arrives late because it travels at the speed of sound. Wolves stalk at dusk, circle the fire at night and back off from a torch. Claude Code split the world between waves of agents, one per domain, and critic agents that never touched the build checked the renders against real photos. 62 seconds of 4K, 1,881 frames, rendered one by one out of the game itself. The wolf is still the weakest thing in it and it doesnt hold 60 fps yet. I ran out of Claude usage before the realism pass finished. Its not even done.
4
35
4,110
Perfect opportunity to spend time with Muse models
Come and build the future of superintelligence at the Meta Global AI Developer Hackathon. Ten days, access to our newest models and $1M in prizes. It’s virtual, worldwide and free to enter. Be the first to know when applications open → bit.ly/4jgQp7e
2
592
Word
When a nerd tells me they don't use agents "because I can write better code than any agent", I usually go "huh, sure buddy, you do you". When they tell they don't use agents for monitoring, I just laugh in their face. Seriously. You can't monitor a production deploy better than an agent can. I dunno about AGI but this specific job is something that an LLM can already do better than every single human in the world. For me, as a systems engineer, having agents monitor my infra changes has been (no joke) more impactful in my day-to-day work than having agents write my code. I started doing this many months ago by asking the agent to monitor explicitly after I triggered a deploy; soon after I moved to using a set of my own Skills. Now our Review Agents team has fully automated this workflow with delightful UX and obviously I'm not going back. It's hard to overstate how powerful this is. All it takes, for me, in production: - The Datadog MCP - The PlanetScale MCP (for all our databases) - Our custom MCP for deploys and internal tooling That's it. Some carefully designed gradual deploy pipelines, and the agent just takes care of everything from there. You can't really deploy SEVs to production anymore. Doing infra like this is pure joy. And the safety is just a tiny part of it. The Rollout agent also does the actually hard part of my job, which is _performance_. When the changes you're deploying are a performance optimization, the agent prepares a monitoring plan that ensures the optimization is correct AND effective. I rarely prompt benchmarks for my optimizations anymore. Just fucking go to prod and the agent is gonna give me the best possible benchmark. If you work on systems or infra for a living, you should be setting up your production environments so you can use Rollouts. Not doing so is actually reckless.
696
Seeing @PalmerLuckey on this brings joy Full circle
this is literally steve jobs coded
2
778
What the magic is Opus 5.5 2 hours of extra high and only 1% weekly usage?
3
13
1,476
It means trying many many things Not tying yourself to one approach Experimenting is cheap so don’t lock yourself to doing things the old way
There are so many seasoned software executives who are completely out of their element with the types of products that will breakout in the next few years. Over the last 20 years, their brains were wired to build apps that deliver value by outputting deterministic results, but AI produces value from probabilistic results. This is such a different way to approach how something should work. Young people who are thinking in this way from the outset will be at a significant advantage -- in the same way they were advantaged when thinking about how products could be social in 2010.
1
2
778
what a throwback track, that alone has me look forward to the UK release of @Muse!
muse is going NATIONWIDE to a television near you this is our first @muse ad ever, going live this weekend! we hope you like it while watching the big game(s)! big ups to @joshginsberg, denise, and the whole meta marketing team ❤️
1
5
1,189
Numman Ali retweeted
This is the future we shall bring into being
Aze Λlter
32,488
40,713
332,687
95,588,413
Performance and price makes me think it’s DeepSeek 4.1 Pro
Here's the DeepSWE result for Union Alpha, a stealth model we just launched. Try it now! Works in every harness. $ ori [code | your-fav-harness] --model=stealth/union-alpha
2
10
1,642
Literally the best paper name Never Give Up This is basically /goal mode for post training
An basic idea in scaling RL: Can we allocate more compute to the harder problems? We did this: If your GRPO group has all wrong completions, sample more with probability P (~0.9) -- in search of more GRPO batches with nonzero gradient. It works! Called "Never Give Up"
4
639
Anyone else seeing very long compactions in Codex? 3m 37s seems pretty excessive to me
2
3
739
“All progress depends on the unreasonable man”
2
444
solid af words right here
the thinking that "everyone will just vibe code their own software" was born from within a bubble what's happening to software development right now has happened many, many times before to other things. you just need to look web 2.0 allowed everyone to write blogs. did everyone write blogs? instagram allowed everyone to share beautiful photos. did everyone share beautiful photos? tiktok allowed everyone to make short videos. did everyone make short videos? now vibe coding allows everyone to build software, and you think everyone will start building software? this time around, people are just different? the mainstream is always a consumer, never a creator. go talk to some people outside of our tech bubble and you'll see - they literally don't give a f about vibe coding
1
813