AI enthusiast & researcher | somi.ai | tinyurl.com/hermes-agent | buildwithclaude.com | AI gets me out of bed, always exploring 🇦🇺

Collabs? DM me ➡️
Pinned Tweet
Introducing Grokbot templates directory on somi.ai A growing collection of viral @bot templates you can use to build your own Grok Bot. • 20+ templates inspired by bots already live on X • Free and open source • Pick a template, customise it, ship your own Steal the patterns. Build something fun.
You can now share templates of your Bots with others.
3
4
3,957
The official Claude and ChatGPT Android apps just lost to a weekend project. Mario built this in 24 hours, on his phone, using the app to build itself. No cloud server. Any model you want. Honestly not surprised. Those apps have always felt like an afterthought to me.
well, this works better than Claude for Android or ChatGPT for Android has ever worked. It runs on my phone directly, no cloud vm. I can live edit it. Its multiplayer. And I can pick any provider/model. Just added artifacts. There is no moat. 24h of slop, developed on the phone, on top of Pi Durable :)
19
Everyone complains AI writing sounds generic. My bigger problem is it's long. Every first draft comes back about twice the length it needs to be. "Be concise" barely dents it, so now I just ask for 120 words and cut from there.
1
1
71
Hakam Kiki sent one guy into One Piece to save the storyline. Ace gets a fire extinguisher. Akainu gets slapped. The Merry gets held together with his bare hands. And everyone gets a forehead kiss.
2
126
Nobody tells this drone to climb. A fruit fly's brain wiring is the pilot. Make the scene drift up, it thinks it's sinking and climbs. Swat at it and the escape neuron fires. The demo's an 850-neuron stand-in and says so. I trust it more for that. github.com/SpikeCalls/FlyDro…
120
Step 12 of 12: lock it by phone. The assembly guide is what got me. Twelve steps in 3D, you can spin every one around, and each nylon washer and servo screw has a label. That's where most home builds stall out. Here it just came with the design.
We are now entering into an era where any product can be created exactly to a consumer's preferences & needs. I vibe-fabricated a dog door with a wifi-controlled lock, perfectly to the specifications & design of my house. I know nothing about metal fabrication or electrical engineering. It's now getting manufactured and delivered in 2 weeks -- for almost the same cost if I bought a mass-produced item off-the-shelf.
1
2
493
16.1 cm vs 16.0 cm. Robot taught by web video vs robot taught by hand-recorded demos. Basically a tie. @runwayml's Praxis-1: 1. That tie, on placement error 2. 0.95 match, simulated vs real tests 3. 3 partners, 3 kinds of robot If I sold robot demo data, I'd be worried.
1
358
46% efficiency. A real watermill gets 60 to 80%. @thebuggeddev built this with Opus 5.5 in one afternoon. Buckets fill, the wheel turns, gears step 6 rpm up to 120, stones grind grain. The 46% is my favorite part. I'd learn more chasing that gap than from a chapter on gears.
136
9 out of 200. That's how many real phishing emails a keyword filter caught. Jev caught 181, at 262 ms each. It did flag 9 normal emails too (the filter flagged 1). Most were real refunds, which honestly look a lot like scams. I'll check 9 refunds to stop 172 more phish.
Decision models for content moderation? I put Jev on phishing 200 real phishing emails from 2025 Our old keyword filter caught 9 Jev caught 181 20x more, at 262 ms an email It also flagged 9 of 200 normal emails (the old filter: 1)
2
349
Two cats, one dance routine, 26 seconds. The orange one in the back is on backup. The fur holds up. The shelves behind them hold up. What gives it away is the legs. That's a dancer under there. No idea which model made it. If you know, I want the name.
305
The wrap-up summary is the least useful thing a coding agent writes. It's usually right. The wrong ones just sound exactly as confident. So I skip it and read the diff.
64
A model that can't write a word just cleared World 1-1. Cloudflare's Clef-flash only picks from options you give it. 248 picks, ~130ms each, local GPU. Clip: @gosrum Most agent steps are multiple choice anyway. That's where I'd use it. huggingface.co/Cloudflare/cl…
1
143
The aircraft manual trick is the one I'm stealing. ASD-STE100 was built for aircraft maintenance docs. About 900 approved words. A 20-word cap on instructions. He ranks explainer videos higher. Not me. Plain text is the only format of the four I can fact-check in a minute.
We'll be spending a lot more time trying to understand the outputs of language models. A few thoughts, tips & tricks: Writing. Something I've had success with: Ask your LLM to explain something in ASD-STE100, it's a controlled language specification originally developed for aerospace maintenance documentation. LLMs well-versed in this language and it comes with heavy constraints on clean writing style that I often find a lot more readable. Sometimes I've tried to soften it a bit e.g. ask for "80% of the way to ASD-STE100" because the spec is quite stringent. But even better: Diagrams / images. Instead of writing, ask your LLM to create a diagram. These can be a lot easier to process, parse, and understand. But even better: Web pages. Ask for output "in HTML" to get a beautiful, interactive webpage. LLMs are getting really good at frontend and can create beautiful experiences, animations, etc. But even better: Explainer videos. The output format I am most bullish on is fully custom / bespoke explainer videos generated on any arbitrary topic. Experiment with things like "Create a 3b1b style video explainer on X. Use my ElevenLabs API key for audio narration". (you'd need an API key for the latter or you can ask your LLM to find you decent free alternatives that use your local compute). This is actually starting to work! In summary: - As LLMs get better, they will do more and more of the legwork autonomously, and a lot more of our work will rise up the abstractions into oversight and understanding. - Luckily, LLMs can help here too because as intelligence and code are increasingly abundant, you can ask for large, custom, discardable software artifacts (e.g. web apps, video explainers) that would have never made sense to create before. Push the boundaries here and you'll be surprised.
1
152
48% of people on a live video call thought Griffin was human. From @tavus: 1) 26 of 54 said "real person" 2) Their old system fooled 1 of 41 3) Doubters caught on in under 20s 4) Every call was one minute long I'd want 10-minute calls before calling it a Turing test pass.
92
A Claude Code mod that stops "rm -rf" and shows you the list first. 9 files, 498 KB. Proceed or cancel. It's called Blast Radius, the second example in this clip. The other two are nice. This is the one I'd install first. Honestly, this one should just ship turned on.
You can now mod Claude Code: - Change how it behaves - Customize the UI - Swap in your own features Write one with a few lines of TypeScript, or have Claude build it for you. Mods ship inside plugins, so you install them with /plugin in the CLI or desktop app. A few examples:
6
962
Four RTX 5090s. 128GB of VRAM. $28,999 assembled. What got me: the whole design is open source. CAD files, parts list, even the BIOS settings. The bare case and riser kit is $1,499. So this video is basically the instruction manual. I'd rather build it.
128 VRAM. 4x RTX 5090. One box serves a whole team. This is what "Build your own AI Datacenter" looks like in real life. autonomous.ai/computer-4 Own your compute. Own your intelligence.
3
201
I'd trade the next model upgrade for one that remembers what I told it on Tuesday. Every morning, same project, explained from scratch. Honestly that costs me more time than any benchmark gap ever has.
1
71
Best use of the MacBook notch I've seen: an Allow button for Claude Code. Coucou's little blob, Mochi, pops out when your agent needs a yes. So it's not stuck in a window you forgot. Open source, by Louis-CFM. Poke it too much and it gets dizzy. github.com/Louis-CFM/coucou
1
5
672
Almost twice as fast in a few hours. Same model, more servers behind it. So if GPT-6.1 Sol felt slow to you yesterday, you were stuck in a traffic jam. I've stopped judging a model's speed in its first few days. Early on, it mostly tells you how bad the traffic is.
GPT-6.1 Sol is our most demanded model pretty much ever both across both the API and subscriptions. Within ChatGPT & Codex, we were under heavy load, but have brough more capacity online and the speed should get much better in the coming hours, reaching almost twice the speed compared to what we served yesterday.
119
Gemini 4 Argon is a lawyer first, coder second. Google's own table: Legal agents: 19.6%. Next best, 6.7%. Finance agents: 65.4% vs 58.9%. DeepSWE: 77.9%, a narrow win. Terminal-bench: last of 4. FrontierSWE: last of 4. Honestly respect that they didn't hide the last places.
2
138