Btw I figured out my own solution to this. I don't like it but it's better than nothing for now
Why did we tie everything to a fucking file system guys
2
716
Just started running our baseline evals for new slate I ran a really untuned harness, astra + bash, on TBench 3 and out of 53 runnable tasks, there were 19 failures due to daytona rate limits lmao Need higher rate limits I guess
5
12
1,075
i have to say, it's pretty cool being one of the top providers on openrouter we serve more kimi k3 tokens than Moonshot itself. pretty wild.
14
8
107
8,636
There is no reason i need to see everything in the terminal. I do not need to see every bash command. This feels silly. Why are we doing this?
1
2
632
akira retweeted
Anthropic's strategy to win back the community: ✅ Support AGENTS.⁠md ✅ Reduce Opus price 40% ✅ Make Opus 5.5 the best agentic coding + writing model ✅ Gave us a banked reset (with no window reset) ❓ Let us use the subscription in any harness? 🤞
187
70
3,141
160,741
My understanding of software engineering has changed dramatically in the past month or so
3
19
2,038
Mutexes are your friend big dawg
5
17
1,438
Being able to focus is a huge multiplier right now and seems like it will only become more important
2
16
686
akira retweeted
accidentally said app instead of domain-specific harness and they kicked me out of SF
107
371
8,787
241,574
Go check out slate programs randomlabs.ai/blog/onyx
Calling it now: the trend of putting control flow in markdown for agents (skills, prompts, etc) is going to look *really* silly in a couple months
1
53
10,093
i just killed instinct. reply for the testflight.
130
11
242
43,139
akira retweeted
if you guys were the inventors of Jev, you’d have invented Jev
40
17
698
37,901
What percent of lab revenue is busted harnesses? It’s double digit %
16
3
90
17,309
There's been a ton of talk about the role of humans in code review, and when and how humans should be signing off on changes. I believe that, long-term, humans have no role in routinely reviewing code.
137
166
1,711
506,563
Ok Slate V2 harness is coming along and it's super cool. Finally doing evals. And I will be posting them. First on the list is just a baseline of astra + bash on tbench
5
26
878
Many such cases I think we shall see a lot of structure output generators that run orders of magnitude faster than LLMs Crazy to see this happen so quickly
We release Needle 3: A Sliceable 8-29MB automation foundation model that can match DeepSeek V4 Flash. One set of weights, every depth from 2 to 20 layers a model of its own, 25-121M parameters at CQ2-bit, built on our Simple Attention Networks and running locally at up to 4k tokens/sec decode speed on a Raspberry Pi 5. Needle does not chat. Every turn is a function call: give it the tools your app exposes and it picks the right ones and fills every argument from what the user said, or hand it a schema and it returns a typed record. Ask for something no tool covers and you get an empty list, not a guess. That trade is lets 121M parameters trained on 360B tokens of structured data beat models 10x their size on mobile tool calls and match 2-3x bigger models on structured JSON extraction. It runs on mobiles, wearables, smart home devices, small robots and microcontrollers, with prebuilt engines for macOS, Linux, Windows, Android, iOS, watchOS, tvOS, the browser and WASI hosts. Try it in your browser: cactuscompute.com/needle
2
3
14
2,267
Has anyone solve agent to agent identity yet?
11
8
1,962
I think they’re buying gpus guys it’s pretty straightforward Buying tokens at spot rates does not require lockup and you need cash flow A big raise makes it possible to buy a cluster for 3 years and do as much as you possibly can with it Good luck
Wow... Instinct in talks to raise at a $10B val, @sequoia + @benchmark looking to lead (@theinformation) 100k+ users, free product, hitting capacity limits + heavy compute costs so they need $$$ Val trajectory: $100M → $500M → $2.5B → ~$10B in a few mos! W/ no revenue... 🤯 Longer term they want to buy their own chips + run their own data centers theinformation.com/articles/…
2
22
5,183
This could be wrong and they could be running the cog playbook but I doubt it
244