Replaced "embedding-drift-monitor" with "telecom-entity-resolution", and upped the repeats to 5. The drift monitor task has an instruction/verifier mismatch due to be fixed in the version of tb.
Here's my first pick for this - tested with Astra Max/Luna, now giving ds4.1 a little run through. Tasks are selected to be indicative, not favouring one model series and simple to run (no GPU/multi-container).
258
Wow! Holy turnaround Anthropic, Opus 5.5 slaps.
Right, and now I'm basically bankrupt. Hits hard when you pay for this by the token from your own wallet. Not a had a session without this pathology yet, not enjoying it any more. Waiting for Opus 5.1 I guess.
1
11
585
Why is the ChatGPT app telling me to download the ChatGPT app?
10
13
593
🤫🤫 openai don't seem to force responses lite through the codex auth route any more🤫🤫
1
2
382
I recently integrated copilot models to fast-agent... The back-end API is genuinely excellent, and offers first-class WebSockets/Responses, Messages and more. Serving is fast and reliable. Easily the best frontier multi-provider experience I've come across.
3
6
286
My approach to new frontier models is to blindly trust them and see how long it takes for regret to sink in.
2
1
22
816
Let's see how cheap we can do a Terminal-Bench 2.1 run with GPT-6-Luna at Flex tier. Give me your score/cost guesses (high reasoning). Wish my credit card luck.
3
248
Here's my first pick for this - tested with Astra Max/Luna, now giving ds4.1 a little run through. Tasks are selected to be indicative, not favouring one model series and simple to run (no GPU/multi-container).
I'm going to start testing harness x model with a subset of tb-4 tasks rather than tb-21. Currently looking at either 18x3 or 20x3 that seem representative and not weighted to one model family. Primary motivation is to keep run cost similar to tb-21 whilst having enough balance to do like-for-like comparison against published leaderboard data. Has anyone else already produced a similar subset..?
4
9
873
Why not call GPT-6-Astra GPT-6 Sol, and what is now GPT-6 Sol GPT-6 Terra? Then there wouldn't be a gap.
3
6
449
I'm going to start testing harness x model with a subset of tb-4 tasks rather than tb-21. Currently looking at either 18x3 or 20x3 that seem representative and not weighted to one model family. Primary motivation is to keep run cost similar to tb-21 whilst having enough balance to do like-for-like comparison against published leaderboard data. Has anyone else already produced a similar subset..?
1
4
1,024
Nice to see Strands approaching fast-agent's efficiency and accuracy. Anyone got any more specifics on the results here? The post itself is a teensy bit opaque. I'm assuming the enclosed bar chart is a %age of 89 trials (as it has a decimal point)...
Today we're excited to share Strands Harness. Strands Harness makes it easier to build reliable, cost-effective, agents with any model. It's open source, too! And equal or better performance to proprietary harnesses with 28% fewer tokens!
6
361
Installed Linux (Omarchy) and now my webcam doesn't work. AMA.
6
10
422
Love this talk by @samuelcolvin - sharing some great stuff on building the best experiences for both agents and humans. Lots of thought has gone in to getting this right.
3
1
11
918
I don't even need to bench Grok 4.7 v. Astra to know it will win (thanks to OAIs safety classifiers). I will of course do it anyway - very keen to compare output token usage.
5
256
AGNTcon/MCPcon Amsterdam was pretty special. Not only because of the great talks and crowds - but because of real excitement about what's coming next for MCP and cool ways that we can wire things together. Invigorating stuff!
6
6
42
1,570
Night out.
2
7
357
Shaun Smith retweeted
don't sleep on @huggingface 's MCP server. it handles the hub like raw storage and compute. agents gets generic bash tools like ls to list, make, append hub repo. agents get jobs tools that can run scripts, commands, containers, servers on any major hardware. best thing is that it plugs into all major agents, both UI and terminal, and is maintained by core MCP maintainer @evalstate hf.co/mcp
11
5
36
2,011
So excited to watch @rachelnabors talk about the death(?) of the browsers. Such cool insights from one of the builders of the web we know today.
1
1
18
536
The GOAT @techgirl1908 opening AGNTcon Amsterdam. Great energy - gonna meet so many cool people over the next couple of days. Come and say Hi!
3
3
26
896