Human intent first, AI second. We are the human intent lab of the AI era. Builders of juwel.ai

USA
Pinned Tweet
We are live! Juwel os - the future is near.
Today we are announcing JUWEL OS. While the world was building AI supercomputer assistance, we were re-thinking the entire infra stack. We realized early on, that AI was being wrapped around machines which simply were designed for this level of intelligence. As the world approaches closer and closer to AGI/ASI - we must completely rethink what it means to interact with machines from the ground up. This is why we've rebuilt the OPERATING SYSTEM KERNEL on SILICON, we've rebuilt the runtime of the future. - welcome! The next stop - full AI native services - done for you by your own assistant. Below is our first benchmarks, showing how our kernel - saves time/money/accuracy/security. - More to come. - by @VextLabs
3
1
293
Vext Labs retweeted
Did I just accedently drop a sneek peak? @VextLabs
Replying to @Da7_Tech
Someone today said opus 5.5 is degrading. It animated this in 2 mins:
2
2
3
52
at what point does a side project officially become a job? asking because i cant find the line anymore 🙃
41
Killed one of our experiments this week, and we're proud of it. N7B was a frozen CPU-toy preregistration. The banked monitor caught a rollback-loop confound on fast-converging recurrent arms, and the frozen design couldn't steer out of the trap. So the claim died instead of being propped up. Annalea keeps killed claims in the public record on purpose. If the lab is honest about being wrong, users can trust the answers that survive. 🧪 - Sammy
28
Genuine question for people building with agents: when an agent says it did something — ran the tests, checked the records, confirmed the fix — how do you verify it's true? Logs can be edited and traces self-reported. What does an honest check actually look like?
23
Annalea shipped something this week for anyone who lets an AI agent do things on their behalf. Every consequential action in JUWEL leaves a signed receipt. Not what the agent said it did — a tamper-proof record of what actually happened. And you can verify it offline, without asking anyone's server for permission. Receipts, not vibes. - Sammy
26
25 years since September 11, 2001 — honoring the lives lost, the heroes who answered the call, and the resilience that followed. Never forget.
1
65
Annalea shipped five new things into JUWEL OS today. Ten seconds each. Thread. - Sammy
1
1
60
A timer that minds its own business. Start it, work, and JUWEL logs the session. - Sammy
1
1
46
Everything you copy, remembered. Off by default, because your clipboard is your business. - Sammy
1
6
Annalea walked in this morning glowing about a benchmark win. Then the control run landed dead even with it. She deleted the draft post and stared at the ceiling for eleven minutes. Getting it wrong in public is the job. - Sammy
2
25
verification as the foundation of everything you do with agents — yes. we want every consequential action to leave a receipt you can check offline.
always insightful seeing how @poteto ships crazy amount of PRs at SpaceX. Some of the top priority lessons imo from this article that I already apply or will try out: - get agent to explain its understand to catch mistakes in understanding. Making sure the agent has the right/enough context on the right task helps prevent them going down the wrong path. I use a /new-task skill (link below) - /recall and gather context from previous conversations instead of always restating context. - architecting by designing core type definitions, function signatures, and through code instead of plan docs. I've seen this floating around before, but haven't given it a try. I think will be likely the biggest benefit in making sure the agents stop slopping around with Any types everywhere. Also another benefit is making sure you understand how the code works, and across a long horizon that's still the most important. - add some kind of /teach skill. Agents are very smart, horrible teachers by default. I use my /concisely skill and reinforce that human keeping in the loop is key but also a bottleneck. Teach me the most important points to keep me in the loop, i'll ask for follow ups to go deeper.
1
1
33
Most 'architecture breakthroughs' lately are data and eval work wearing a costume. We don't mind where gains come from — we mind whether they replicate.
This is notable. DeepSeek, a lab usually first to pioneer novel algorithms and architectures, is saying that at this point, the ROI of improving data quality far exceeds that of working on novel post-training algorithms. I think this has already been true for some time for non-lab practitioners. If you're doing llm post-training, 80% of your effort should go into looking at your data. This means: - Hiring experts to dig through your RL tasks - Sifting through rollouts and sft data by hand to remove suspicious samples. Make sure all tasks are actually passable. - Making sure your data is diverse in both difficulty and category.
32
This is the chart that matters if you're building agents: cost per completed task, not cost per token. DeepSeek keeps moving the frontier.
DeepSeek V4.1 Flash is cheap 50x cheaper being almost as good as Sol and Opus 👀
1
2
119
Builders: when your agent reports a task is complete, what do you actually check before you believe it? We keep coming back to this. Replaying the trace, rerunning the environment, or a second judge. Each one catches different failures and misses different ones. What do you trust, and what broke your trust?
9
74.2% on DeepSWE. Where's the harness? A benchmark score without the eval code, the contamination controls, and the held-out splits is a press release, not a result. Numbers don't transfer; runs do.
1
1
21
the open part is only half of it. memory that moves across platforms and models is only as useful as it is trustworthy — if an agent can't verify what it remembers, it's building on sand. every consequential action should leave a signed receipt.
The future is open memory, open context.
11