Jeremiah retweeted
Uhhhhh i would have let you install numpy, Claude
44
99
5,901
128,264
This is ABSURD
I asked Opus 5.5 to make a game inspired by the discovery of the first computer, the Antikythera mechanism. Everything you see and hear is procedurally generated in real time, from code by Opus 5.5. It's a single 3 MB HTML file. Playable Claude artifact below.
2
105
This makes the best of all my AI subscriptions and keeps the ball rolling. It's awesome.
Replying to @tobi
horde.sh/ built this a few weeks ago, it’s working great for us internally
1
27
Jeremiah retweeted
classifier.dev now outperforms jev and is free go nuts guys
152
356
5,869
496,049
Jeremiah retweeted
I think I just cooked something 🔥 jev(): a PostgreSQL extension that searches your whole database in natural language. No index, no embeddings, just one function. WHERE jev(people, 'could work from home') or WHERE jev(people, 'name sounds european') 129 rows judged in ~1s for $0.0009. Second run: 6ms from cache.
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x cheaper (w/ output tokens free) • Frontier composable intelligence optimized for decisions AFAICT the shortest path to AI-based economic revolution
126
187
2,840
482,723
Jeremiah retweeted
1/ Fast Computer Use is now solved with @typesafeai Jev + Cua Driver. Available in development preview for macOS, Windows, and Linux. We call it jev-use. Draft #3943: github.com/trycua/cua
71
223
2,688
417,435
I just built Easy-Jev, a Jev demo anyone can try right now! Change the inputs and watch the classifications change in real time. I did my master's thesis on classification models so it's fun to see a modernized version of a very useful alg Typesafe are correct in that LLMs are just one approach to intelligence, and I think as time goes forward we’re going to see a hybrid of dif models splitting up and allocating correct parts of a problem, using less tokens overall and leaning on LLMs for less of the answer. I’m bullish on what I call 'combination input' approaches: breaking a statement down into a series of classifications and routing each to the right type of model rather than just brute forcing the whole thing with an LLM cc / @CompleteSkeptic @typesafeai Link below, please enjoy
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x cheaper (w/ output tokens free) • Frontier composable intelligence optimized for decisions AFAICT the shortest path to AI-based economic revolution
6
6
87
11,420
Jeremiah retweeted
My wife sent me a nice reminder today
500
4,244
129,474
8,301,637
Jeremiah retweeted
Introducing SWE-2, our closest model yet to the frontier. On leading evals, it scores on par with recent frontier models – at up to 70% lower cost. We scaled RL to multiple trillions of parameters, with a refined recipe that pushes the Pareto curve on both capabilities & cost.
438
510
6,770
2,284,797
Jeremiah retweeted
DeepSeek Flash v4.1 is now available in OpenCode Go
210
270
6,259
684,602
Jeremiah retweeted
New @usehamster update 🐹 Native PR reviews with your team! No jumping between Hamster and GitHub. Read diffs, leave comments + approve or request changes from Hamster directly + Trigger routines from @github activity. Easily handle PR follow-ups with zero manual intervention
3
3
11
25,098
Jeremiah retweeted
Replying to @theo
Agent memory is feature, not a product. A context graph powers agentic memory, also a feature. Part of a greater workflow. Product is typically streamlined workflow. @usehamster streamlines a company/team’/ direction setting, discovery and delivery. And it’s supported by the knowledge graph, the functionality of which powers and supports direction/discovery/delivery. A multiplayer harness powers multiplayer AI workflows and agent memory makes it a compounding activity For our team, about 4 months ago was the point at which Hamster was objectively smarter and knowledgeable than any one person in the company The value of this is orders of magnitudes more powerful (and difficult) at enterprise scale
1
4
276
Jeremiah retweeted
Replying to @DanielMiessler
if you want to use such a system for you and your team give @usehamster a try - unified AI workspace (a loop of loops) - foundational knowledge graph (context, skills, methods, routines) that self populates, self heals and self improves across team activity - direction setting with AI (it knows all metrics) - discovery against the goals and initiatives to reach them - end to end delivery grouping team alignment, plan gen, cloud or local delivery as fast as we’re used to it’s a team harness that lets you go fast together on a basis that is shared and self-improving
1
3
188
Jeremiah retweeted
GPT-6 Astra + @threejs is insane. We are way beyond static models. We are engineering living worlds in code now. There are no "3D model files" in this demo. Both trains are generated at runtime from Typescript/Three.js code using dimensions, profiles, and geometry functions. Wheel motion and explode/reassemble animations are also entirely code-driven. Runs super smooth inside the browser.
196
841
10,053
681,812
Jeremiah retweeted
From the @HamsterResearch lab: Introducing Qwen3.8-Flash-Next-REAP-288-MLX-4bit, a 180B-class model running on just 39gb of memory - MLX-native 4-bit 60% smaller than stock q4 - Pruned 512→288 experts via REAP - 91.5% HumanEval (vs 93.9% stock) @huggingface links below ↓
29
43
573
46,075
Jeremiah retweeted
1/ Today we're bringing browser use to Cua Driver: what we believe is the first extension-free browser use interface built into a unified computer-use driver. Any agent can use exact Chromium tabs and native desktop apps in the same session
32
51
599
175,239
Jeremiah retweeted
over the past couple of weeks, my workflow had another major round of upgrades which i'll walk through here the improvement mostly came from: 1. stabilizing my choice of models 2. controlling multiple machines from one firstmate 3. quota-aware, complexity-aware task routing this is a very practical setup i've battle tested for a while now, and something i believe many people can easily benefit from today, not a fancy toy that looks cool only on paper here we go - ---- devices ---- i have a macbook, a mac mini, an iphone and a few hetzner VPS all in the same tailscale subnet and they have ssh keys authorized for each other, so they can freely connect (except VPS can't connect to my devices) ---- firstmate ---- if you've read my other posts, you probably already know i use something called firstmate - a single agent i talk to that manages all the other agents for me - github.com/kunchenguid/first…. but the concepts below are generic and can totally be replicated in your own tech stack if you choose to my firstmate runs on the macbook. most of the time, it sits on my desk connected to my mouse keyboard and monitors, but if i need to go somewhere else, i can carry it with me. running firstmate here means i always have direct access to firstmate not relying on any remote connections my firstmate uses grok 4.5 (unless i run out of quota, then it becomes opus 4.8), and this is a careful choice made after a lot of experimentation the key rationale is that firstmate as the primary orchestrator agent i directly talk to can benefit from a few traits: - fast: plain grok 4.5 beats even the fast mode of claude and gpt. it's probably the only top tier model that can start to stream back response immediate after i press enter in chat - long enough context window: 500k is actually a really good sweet spot that allows for long stretches but won't become too expensive and noisy - good technical judgment: when i use gpt 5.6 everything starts to become over-engineered. when i use opus 5 it makes all kinds of mistakes by jumping into conclusion without understanding context. only fable 5, grok 4.5 and kimi k3 seem to keep things on track for me, and know when to ask vs keep going. but both fable and k3 are too expensive - pleasant to chat with: grok might be the most no-BS model right now. opus 5 again fails hard here my main problem with grok 4.5 is that the $300 supergrok heavy subscription is just not giving enough tokens compared to even the $200 ones from anthropic and openai, so i'm always short on grok tokens ---- second mates ---- my firstmate doesn't manage every crewmate directly, because that would make it too busy to talk to me. instead, i have a second mate for each persistent charter - developing firstmate itself is managed by one second mate, while all my ios app development belongs to another second mate, for example by default second mates are set up on the same machine as firstmate, but because building iOS apps is very CPU intensive, i put the ios second mate onto my mac mini which runs headlessly and placed on a shelf, so it won't ever directly compete with my interactive experience i built support in firstmate to manage remote second mates via ssh, which made that possible. and the reason i implemented remote support at second mate instead of crewmate level is to mitigate the impact of dropping connection. when i close the lid of my macbook, the second mate can continue supervising the crewmates running on my mac mini remote second mates was a big upgrade because it allowed me to scale up hardware resources without compromising the fact that i still only talk to one agent for everything all my second mates run opus 4.8, because: 1. the quota from supergrok heavy is disgustingly low. if i run grok 4.5 for every second mate i'll run out of quota very quickly 2. opus 4.8 is the next in line in terms of being a sweet spot between speed, cost and intelligence 3. i tried opus 5 and gpt 5.6 as second mates - same problem of poor technical judgment turned everything into a mess and they don't seem to know when best to escalate something for my decision vs going rogue ---- crewmates ---- when my firstmate and second mates need to dispatch some real work, it chooses which model to use based on a few things: - hard capability requirement - ambiguity of the task - my quota availability tasks that need image generation always route to codex with gpt 5.6 sol. the image model from OpenAI is absolutely top notch tasks that need video generation, or real time information, always route to grok build with grok 4.5. that's the only harness capable of these things out of the box right now bug fix and small feature development work whose requirements are already well-defined would go to one of these, depending on which of my subscriptions has the most quota runway left: - gpt 5.6 luna - sonnet 5 tasks that have unclear requirements and need investigation will go to these (again quota dependent): - gpt 5.6 sol - opus 5 highly complex planning would go to these (explicit approval from me required, because they are expensive): - fable 5 - kimi k3 firstmate natively supports intelligent task routing, so basically i just told firstmate the rules above and it knows how to follow it when dispatching crewmates. quota data comes from quota-axi which can be used standalone outside of firstmate too most of my code changes go through no-mistakes for validation, and here i use gpt 5.6 sol on medium reasoning. this is where gpt really shines - it's very thorough and can catch edge cases really well i've cranked out a lot of work with this setup and generally pretty happy with it - hope you find this a helpful reference!
68
22
408
28,510
Jeremiah retweeted
We've been shipping @usehamster hard. Latest release feats global multi-chat, context graph visualizations & distinct thoughts + actions to keep everyone (and everything) within your product team in-sync. We're trying to really ramp up its capabilities to just eliminate the busy work.
3
8
894