Cofounder, AI @EQTYLab / prev dc - complex systems / veteran / For fun: @plannotator

CA
Anthropic code review this, clanker review that ... why don't you shut up and review+annotate your own code.... (yes im a loser who still manually reviews code) Originally inspired by a bunch of feature requests and then seeing @dillon_mulroy tweet a similar cool ux. @plannotator for reviewing plans (primary focus) and code, fully oss. OpenCode, @badlogicgames 's pi.dev, and Claude Code and other clankers
18
17
316
73,876
Michael Ramos retweeted
show me your custom cloudflare dashboards for birthday week ❤️‍🔥
5
8
47
3,392
Sharing some prompt injection "jailbreak" benchmarks I've been running with Jev, competing against a bunch of standard guardrail classifiers including Meta's. Part of it is a standard open suite that's getting dated. Jev wins that one outright, and wins the newest attack set too. It loses on the older sets and can't run at a tight false-alarm budget. It failed some of our internal benchmarks when something more nefarious is going on & is obscured, but there's really not any good off the shelf classifier that could be used for those type of scenarios. You can see a glimpse of this with how every model performed poorly at allenai's WildJailbreak. Since the poker eval I've been trying to pin down where Jev fits. My hunch is perfect-information calls with little strategy in them. Still testing it.
9
1
25
1,295
We are in the age of orchestration. It is very effective to work through 1 agent per project, while it delegates tasks to durable subagents. Hence Cursor/Claude Code Projects (not that they offer anything you cant do today without them). The approach doesn't require alchemy or some stupid metaphor philosophy approach. Just a proper subagents implementation, and integrated-monitoring. Some popular harnesses fuck these things up. Here's 2 extensions to get it right in Pi: 1. subagents github.com/nicobailon/pi-sub… 2. Monitors/polling/loops github.com/joelhooks/pi-unti…
22
6
169
13,895
Codex needs actual monitors. Scheduled reminders as implemented eats a whole turn of tokens every time. It is not viable for any type of polling situation. Instead I need to set a programmatic monitor that only alerts the agent when needed. And we all need MCP events/triggers/tasks/whatever asap
10
39
2,462
Michael Ramos retweeted
The latest Plannotator is out @MermaidChart's 12.0 release is awesome so we've shipped it plus theming sync and new interaction capabilities.
4
7
128
5,920
If you are ab testing engineer workflows using your product - you do not care about engineers.
Replying to @Steve_Yegge
Yep you're part of the experiment. Yesterday I found out Anthropic was quietly running A/B tests with my system prompts in Claude Code without my knowledge 😢 lnkd.in/p/gahaK-Uh
7
1
25
1,839
Jev is a 🐠 / not good at poker. Up to you to decide what intelligent and consequential decisions you're willing to have this thing make for you.
4
27
2,699
writeup:
Here is my writeup on jev & making decisions in a game of poker backnotprop.com/blog/jev-pok…
167
Just spotted this frontier lab billboard in SF
3
2
14
877
This is the future in that we really only need to interface with one agent. Age of orchestration turns into us talking to a single assistant. Easier on us cognitively. But you could already do this very well in CC/subagents and I hope the core does not degrade. I just need a simple agent interface and high class sub agents implementation.
Projects now run from one conversation, starting in Claude Code. You describe what needs doing, and Claude directs parallel threads that keep working after you close your laptop. In beta today for select Pro and Max users in cloud sessions; coming to all Claude users soon.
2
22
2,122
Haiku 4.5 with reasoning turned off
1
1
380
Poker eval was iterated on quite a bit. At one point I Showed the opponent's hand, yet Jev still made a poor decision. Jev would respond correctly when I gave it very clear the state of the world, like when it was in a non-favorable position like "Opponent has nut flush. Hero has 0 outs" And you could say well you need to tell the model versus make it derive from seeing the opponent's hand, but in my mind this is all a wash - because half the showcases on our X are about, you know, can this model make intelligent decisions? So I don't think I'd use it to classify insecure code or make decisions about how to navigate across a browser for any consequential task. I do have plenty of fun use cases in mind though.
1
5
300
Michael Ramos retweeted
Today we introduce TIN: a powerful and reliable full-text search extension for Postgres. TIN works with complicated WHERE clauses, replication, backups, and maintains correct transaction visibility. It's also mind-blowingly fast.
58
114
1,977
501,372
Michael Ramos retweeted
Replying to @backnotprop
Thanks for mentioning GLiCLass. We initially designed it for information extraction workflows. But actually, the long-term strategy was to develop systems like Jev, we also introduced RL-based training for such models in our paper: arxiv.org/pdf/2508.07662 Also, please check our newer, more generalist model: huggingface.co/knowledgator/…
1
5
111
I deleted a tweet that was starting to go viral. I misspoke on the benchmark comparison and it would have been wrong to spread that. Still worth knowing about GLiClass... github.com/knowledgator/glic… huggingface.co/models?search… There are open models for zero-shot performant classification. Maybe useful as a foundation for Jev-like systems. What Jev has done is impressive in terms of the RL and global calibration.
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x cheaper (w/ output tokens free) • Frontier composable intelligence optimized for decisions AFAICT the shortest path to AI-based economic revolution
3
1
30
2,794
ok ok, getting attention... I mispoke on the benchmark comparison. GLiClass and Jev are not being measured on the same tasks or metrics, so the similar-looking numbers do not establish similar performance. Jev has some cool advantages-RL/calibration, funding/team. GLiClass is an open foundation for building Jev-like systems. But that is a different claim I should have spoke better about. My bad on that part.
1
4
257