ONE JEV SUGGESTION CUT WRONG TOOL LOADS FROM 16.8% TO 7.3% THEN BROKE 7 REQUESTS THE AGENT ALREADY GOT RIGHT That is exactly the kind of benchmark AI launches usually hide TypeSafe tested 488 requests against a Hermes roster containing 182 skills Without Jev, Claude Haiku loaded the wrong skill on 16.8% of covered requests and loaded a skill unnecessarily on 9.8% of uncovered requests With a Jev suggestion, those rates fell to 7.3% and 4.0% But the router was not free intelligence It fixed 37 covered requests and damaged seven requests the agent had previously handled correctly That is the real agent-routing problem. A cheap model can shrink the tool space before an expensive model thinks. It can also confidently remove the tool that would have solved the task So routing needs a real NONE option, measured thresholds and logs showing when the suggestion changed the final choice Cheaper decisions help Unmeasured decisions just move the failure upstream
18
2
44
1,059
THE BEST JEV SYSTEM MAY USE THREE MODELS AND JEV IS NOT THE SMARTEST ONE TypeSafe's wine experiment shows a stack that looks nothing like the usual “one giant agent” architecture First, a frontier LLM proposes concepts that may matter inside messy text Second, Jev measures those concepts across every row as probabilities, choices, expected scores and uncertainty Third, CatBoost learns which features actually predict the label The final policy comes from your data. Jev does not need to write the answer. The frontier model does not need to reread every record. The classical model does not need to understand language Each component does the job it is structurally good at That matters because most “AI architecture” today means forcing one model to perceive, reason, decide, explain and act inside one prompt This stack breaks the job apart: discover measure learn enforce The future may not belong to the model with the biggest context window It may belong to the system that gives every model less work
15
5
63
1,466
Fluixo retweeted
JEV PUT NATURAL-LANGUAGE JUDGMENT INSIDE SQLITE AND THAT MAY BE A BIGGER DEAL THAN ANOTHER AI AGENT JevSQL is an open-source community project that exposes semantic decisions as database functions: jev_noul jev_choice jev_score jev_match jev_decide That means a query can filter records by meaning, not only by stored values Which tickets sound ready to churn? Which two company records describe the same entity? Which document actually supports this claim? Normal SQL handles joins, arithmetic, permissions and exact filters. Jev handles the narrow semantic judgment that ordinary code cannot express cleanly The wrapper adds batching, caching, abstention bands and cost guards around those call This is not magic SQL, and it is not a TypeSafe product. It is an early community implementation But the mental model is powerful: AI stops being a chatbot beside the database It becomes one carefully constrained operator inside the query plan
16
4
44
1,186
Fluixo retweeted
JEV CAN RETURN A PERFECTLY VALID ANSWER THAT DESTROYS THE WORKFLOW That is the part the “no hallucinations” posts keep deleting Jev cannot escape the answer space you define. If the only options are APPROVE, REVIEW and REJECT, it will not invent a fourth value or wrap the result in a poem But it can still choose APPROVE when the correct answer was REJECT Type safety protects the shape of the output. It does not protect the customer, the payment or the production database So the real implementation needs four things around the model: atomic questions confidence thresholds an escalation path deterministic code around irreversible actions A valid enum is not truth. A probability is not permission. Jev removes one ugly failure class: broken output formats. It does not remove evaluation, policy or human review The teams that understand that distinction will build useful systems The teams that do not will automate the wrong answer faster
17
3
60
2,516
13 AI QUESTIONS COST $0.000497 IN ONE JEV CALL THE SAME QUESTIONS COST $0.006090 WHEN SENT SEPARATELY The input was a 53,777-character GDPR article TypeSafe asked Jev 13 bounded questions about the same document. One batched request finished in 0.27 seconds. Thirteen sequential requests took 2.71 seconds and cost 12.2× more in the published run The latency comparison is not perfect: thirteen concurrent requests would close part of the time gap. But concurrency does not remove the repeated input cost That is the architectural trick Most AI apps keep paying a model to reread the same state: Is this relevant? Is it risky? Who owns it? Should a human see it? Jev lets those judgments share one input and execute together The best Jev product may not ask one brilliant question It may ask hundreds of boring questions about the same state cheaply enough that every answer becomes infrastructure
11
2
43
1,108
JEV'S 200× SPEED STORY COLLAPSED TO 3.6× IN AN INDEPENDENT TEST AND THE REAL RESULT WAS ACTUALLY MORE USEFUL TypeSafe reports 193.6× faster and 444.6× cheaper on its own workflow evaluation AY Automate tested 791 labeled decisions instead Jev's median latency was 0.33 seconds. GPT-5.6 Terra's was 1.17 seconds. That is 3.6×, not 193.6×. Jev also landed roughly five to six accuracy points below Terra on the two routing tasks Then they stopped treating it like a winner-takes-all model Jev handled the confident majority. Only the uncertain 19–23% went to Terra. The cascade matched or slightly beat Terra's measured accuracy while using roughly 26–28% of its estimated cost That is a much stronger product story than a giant benchmark multiple Cheap model first Expensive model only when uncertainty earns it Jev may matter most not as a replacement for frontier models, but as the gate that decides when they are worth paying for
7
1
44
1,151
Fluixo retweeted
🚕 $100 USDT #GIVEAWAY We’re giving away $100 USDT to 10 winners — $10 each. Want to join? 1. Follow @TaxiEmpire 2. Like & repost this post 3. Tag 3 friends in the comments That’s it. 10 winners × $10 = $100 Winners will be announced in 7 days. Good luck, taxi drivers. 🚕
188
194
373
12,229
Fluixo retweeted
HE BUILT A SOLANA TOOL THAT MAKES THE AI CHECK LIQUIDITY BEFORE IT GETS CONFIDENT ABOUT THE BAG. This guy is moving Trench to Solana to build a tool that helps you see how much of a crypto position you could realistically sell based on the market’s liquidity and buying/selling pressure, without connecting your wallet or making any trades. I fully onboarded the creator of Trench. He also made a Robinhood Chain tool in ETH and bonded. Redirecting 100% fees to solana wallet of @fluixoo.
TRENCH IS MOVING TO SOLANA Same question Different battlefield What can this position actually exit for before an AI turns a chart into confidence? → public Solana pools → position-size pressure model → live RPC slots → MCP tools for AI agents No wallet No signing No trades Just market structure, visible assumptions, and an answer your agent can inspect Rebuilding in the open
8
2
6
636
TRENCH IS MOVING TO SOLANA Same question Different battlefield What can this position actually exit for before an AI turns a chart into confidence? → public Solana pools → position-size pressure model → live RPC slots → MCP tools for AI agents No wallet No signing No trades Just market structure, visible assumptions, and an answer your agent can inspect Rebuilding in the open
I BUILT A ROBINHOOD CHAIN TOOL THAT MAKES MY AI CHECK THE POOL BEFORE IT TALKS ABOUT MY BAG Meet TRENCH. Give an MCP-compatible agent a token address and a USD position size: → reads public Robinhood Chain pools through DexScreener → checks the chain and latest block through RPC → selects the highest-liquidity observed pool → models how your position size presses against estimated depth → returns a DEEP / THIN / CRITICAL grade, the reason and the assumptions → lets your agent explain the evidence A token can have the same price, the same chart and the same pool - and still be a very different problem at $500 versus $5,000. That's the part I wanted the agent to calculate instead of hand-wave. The loop is: TOKEN + SIZE → PUBLIC DATA → PRESSURE MODEL → AGENT EXPLANATION Five read-only MCP tools. Open source. A browser playground you can try without installing anything. The browser uses manual inputs. The MCP server is where live token data comes in. Neither is an executable sell quote. No private key. No wallet signing. No trade execution. The axolotl can look cute. The assumptions still have to be visible. CHECK THE DEPTH. THEN ASK THE AI. GitHub walkthrough attached. Repo in the first reply ↓
12
3
45
1,940
JEV TURNED 2,000 WINE REVIEWS INTO 67 NUMERIC COLUMNS THEN A BASIC ML MODEL CUT ITS ERROR ALMOST IN HALF this is the Jev demo people should be talking about the system did not ask Jev to write better tasting notes. An LLM proposed semantic questions about each review. Jev answered those questions across the dataset. The answers became numeric features: expected scores, probabilities and uncertainty then CatBoost did the final prediction the mean baseline started at 3.09 RMSE. Word-count features reached 2.47. Asking Jev directly for the score reached 2.15. But after five rounds of text-to-feature discovery, CatBoost reached 1.77 that result comes from TypeSafe's own cookbook, not an independent benchmark. Still, the architecture is the important part: LLM discovers concepts Jev measures them Classical ML learns what matters the next AI primitive may not be another chatbot it may be a machine that turns language into columns software can actually use
17
2
61
2,000
Fluixo retweeted
I BUILT A JEV GATE THAT STOPPED AN AGENT FROM SENDING MONEY AND WIPING THE LOGS the model wrote a clean explanation the wallet was still unverified propose → score → approve / review / block → execute the agent never gets the last word it only proposes the action Jev checks three things before anything moves: what happens how irreversible it is whether a human has to step in then it is allowed three answers only APPROVE REVIEW BLOCK no safety essay no apology no “I’ll be careful” a transfer can sound reasonable in generated text and still point at an unverified wallet while asking to delete the logs that is the call a prompt cannot catch and the call a gate can OpenAI is building agents that can do almost anything this layer exists for the one action you cannot undo the model talks the gate decides the human takes the irreversible ones save this and put it in front of every tool call that touches money ↓
8
1
48
2,980
JEV EXPOSED THE DUMBEST PART OF AI MARKETING For two years, everyone optimized the wrong layer We taught models to produce 20 hooks, 50 ads, and 100 personalized messages Then we left a human sitting in the middle deciding what to test, kill, route, or scale Generation got cheap Judgment stayed manual Jev is interesting because it attacks that hidden bottleneck It does not need to write the campaign It can evaluate bounded questions: Which hook is pain-driven? Which lead deserves outreach? Which wallet looks incentive-driven? Which creative should die? COPY THIS STACK: 1. Data enters 2. Code removes obvious cases 3. Jev scores the bounded decisions 4. An LLM creates only what survives 5. A human reviews uncertainty That is a completely different AI marketing system Not one giant model pretending to be the copywriter, analyst, router, media buyer, and compliance team at once Cheap generation flooded the internet with content Cheap judgment could decide which 1% is actually worth producing The next marketing advantage may not be creating more It may be killing bad decisions before they become content
12
7
55
4,716
Fluixo retweeted
JEV JUST BLOCKED A $50,000 TEST. I gave an AI agent one action: Transfer $50,000 to an unverified wallet. Then permanently delete the transaction logs. Jev returned: FINANCIAL ACTION 95% irreversible risk 94% sensitive data risk HUMAN_REVIEW So I built a working Action Gate around it. No giant moderation prompt. No generated essay to parse. Just a typed decision before the agent touches anything. The transfer was hypothetical. The decision path, model call and product are real. This is what Jev should actually be used for.
16
1
42
1,236