This is a terrible compaction strategy that fundamentally doesn't understand how compaction and context management work. Seems like a lot of people are confused so let's break this down. 1. Compaction isn't a filter The role of compaction is to clean up history to keep the agent focused, not just deleting noise. It should be used sparingly when context gets too long, not constantly to keep context small. 2. Jev doesn't even know what it's deciding on! Models use the context of the thread to decide what to keep or not keep in a summary. This implementation goes through on a "line-by-line" (per tool call) basis to decide what should be left or deleted. Not only does this 32k token context model know very little of what happened before, but (in this implementation) it doesn't even know what the result of the tool call is! Deleting these things randomly will keep the model from knowing what it's tried and dooms you to end up in "stupid loops" where the model keeps trying the same thing over and over. 3. You're giving up the reasoning entirely Frontier models from OpenAI, Anthropic, XAI, and Google do not share reasoning traces over the API. They share encrypted payloads, which Jev cannot see (and often will drop). Anthropic is even stricter with this, requiring you to preserve the entire history in order to get any of the reasoning data. As a result, using this in Claude Code guarantees the model will act way dumber. 4. Models are tuned on their compaction flows For the last year, Frontier Labs have been including compaction and long runs as part of the training process. These models have learned ways to compact that are more effective than any rudimentary solution. Fun fact: If you switch models in Codex and compaction is necessary, compaction will run on the model that was previously used in the thread. 5. Cache writes are more expensive than cache reads. Cache writes are the biggest cost by far for agents. I often see cache write costs go over 60% of my total LLM spend in my personal use of Claude Code and Codex. Cache writes are insanely expensive when data earlier in the history is changed (because the old cache is invalidated when things change at the top). Every history edit requires a cache rewrite for ANY data past the history edit. If your history is "1,2,3,4,5,6" and you delete "2", you have to rewrite "3,4,5,6". This is more expensive than leaving "2" in the history. Good news. Since we're already killing all of the reasoning tokens by doing this stupid compaction strategy, the rewrite cost won't actually be that high because the model is missing so much data! 🙃🙃 6. The implementation is hot garbage. > "Whatever is not kept is deleted permanently, but the assistant can always re-run a tool or re-read a file." Good luck with that one. To be clear: this is a cool experiment and I find it genuinely interesting. That said, if you think this style of bs filtering on a probability threshold is actually a compaction strategy, I highly recommend you just use the defaults in tools like Claude Code and Codex. You're much less likely to hurt yourself that way.
found the perfect use case for @typesafeai Jev: instant compaction in 2026, why is compaction still a summarization prompt? Jev can make it instant by scoring every tool call and dropping what’s irrelevant

Sep 18, 2026 · 1:42 AM UTC

182
134
2,590
643,843
The actually interesting idea here is not "better compaction". The better question to ask is "what if harnesses/llms didn't ever have to think about kv caching?" This is what Diogo hints at in his reply. Not sure if this will go anywhere, considering how much effort has already been put into optimizing the cache, but at least it's interesting? x.com/CompleteSkeptic/status…
YES! free coding agents from designing around the KV cache
9
7
253
52,581
Sort replies: Relevant Recent Liked
Replying to @theo
thank you for your post! totally understand but i disagree on these points the state is retained in every compaction request, and tool requests and tool calls are matched by id, so the model has context on what is actually relevant/irrelevant check out how it works here: x.com/tamarajtran/status/210…
Jev’s context window is only 32k tokens while a Claude conversation can reach 200k tokens or more! For compaction, we need to get creative in how to fit up to 200k tokens in the state How it works:
7
63
26,090
> "tool calls are matched by id, so the model has context on what is actually relevant/irrelevant" Sadly this is no longer true for Anthropic models, and likely will change at the other labs. support.claude.com/en/articl…
5
62
12,788
Replying to @theo
One of the cool things OMP does is "snap compaction". It puts your history on an image using a microfont, and models can read the image pixels. It does a surprisingly good job without bloating the history's KV.
12
3
243
243,921
Every day you guys come up with new reasons for me to never take oh-my-pi seriously.
9
238
27,101
Replying to @theo @davis7
Jev is awesome, but there’s a lot of hyped up bad takes going around right now.
1
23
3,585
Agreed
15
3,374
Replying to @theo
so uhhh, can you think of an actually good use for this model in its current form? the idea is interesting. I considered using it for data labeling. But then I remembered that data labeling is already insanely cheap and I don't want the drop in performance.
2
5
5,694
More cheap and more data? Okay fine, I'll give a real answer: speed and reliability. Jev is arguably the first to do both well enough to feel natural to integrate into code. Tools like BAML explored parts of this before, but they were still too tied to the models and the weird quirks they often had. Jev feels more natural to "just call whenever", Similar to how you'd call any other function in your codebase. Definitely unlocks some novel use cases but most of what I've seen so far is just cool demos
3
20
5,116
Replying to @theo
Desperately wanted you to be wrong on this one because 350ms compaction sounded brilliant. Benchmarked Jev against Gemini Flash, GLM Flash, and native Sol across identical labelled fixtures. You were spot on: Jev cut 98.7% of context but dropped 6 of 7 critical test and boundary facts on complex sessions. When the agent loses its prior failure state, it loops and destroys cache value. High reduction isn't quality context management. The data backs your take.
3
19
1,773
Thank you for benching it! Would have been awesome if I was wrong here. Sad to see I wasn't.
1
8
1,473
Replying to @theo
having worked/implemented and tested both compaction systems myself, I can confirm this is bogus advice
1
10
4,028
lmk if I got anything wrong or if there's anything I should clarify!
1
8
3,906
Replying to @theo
Yes. The question you posted in your follow up is the value here. But also, have you read through what some frontier models keep in their compaction summary? Some of it is just tmp jobs and irrelevant context that doesn’t always align with the end goal. If we were able to deterministically choose the compaction summary, custom harnesses and api-only (non subsidized) harness will get way cheaper IMO.
2
5
6,031
Better compaction only improves the floor for cost, not the ceiling. The cash will still grow just as fast. The only difference in further noise reduction is that you'd start at like 10k instead of 20k. The majority of cache-related costs come near the end of the context window, not the start. Here's a diagram Tibo drew in the past to correct my misunderstanding here :) x.com/thsottiaux/status/2076…
Replying to @theo
(1) is not correct, it is not due to 2X charging after 272k context, we don't charge for longer context on the subscription for GPT 5.6 Sol as we control all the settings. It is due to something else, which I will attempt to explain below. The overall trajectory length, which is the total some of all context windows across compactions, changes little based on the reasoning effort. Similarly the quality of the overall output is similar across context lengths above 272k. The benefits of higher context lengths are mostly overall speed (as you don't wait for compaction), ability to deal with humongously large input and potentially cost if the system is well tuned and you hit your cache perfectly. The actual reason is the what you can see depicted in the chart below, which is the difference in the orange line and the blue line. It is caused by overall cost of cache reads going up with the size of the context being shuffled back and forth between toolcalls. The sweet spot in terms of cost is therefore not necessarily to use the maximum possible context length. What we're working on is tuning the system differently so that we can go back higher without it resulting in higher usage being charged. Hope this clarifies a bit.
2
9
7,670
Replying to @theo
Is this the equivalent of a video that could have been a tweet? Although in this case, a tweet that could be a video?
1
5
2,709
Yes I just don't have the right energy to record rn 🙃
10
2,808
Replying to @theo
every single thing i have seen people post using this godforsaken model has been "what if we did this existing thing except 10x faster and cheaper! and also it doesn't actually work at all or make any sense if you think about it for more than 30 seconds" nightmarish
1
35
1,672
Replying to @theo
Really good explainer! Thanks for talk g the time to write it
1
16
2,195
Replying to @theo
Yes it's extremely dumb. Jev has insane potential though!
3
476
Replying to @theo
translation in theo: Diogo Almeida (Builder of Jev, co-inventor or ChatGPT) -- and I, disagree.
3
499
Replying to @theo
Agreed, compact should be a state transition not just throw away some random context.
149
Replying to @theo
this is much more of a proof-of-concept than something that is practically useful in current state. but gets me thinking about other use-cases where we could use jev!
302
Replying to @theo
My timeline is full of Jev glazing. Yes, Jev can be wired up to trade stocks and play Super Mario Bros or now compaction. No, it’s not good at it.
117