Science background. AI foreground. probably broke prod again. I create things I use the most myself.

localhost
Steve ๐Ÿ‡บ๐Ÿ‡ธ retweeted
Waitlist open. Apply: agnt.trade Close in 24 hours.
3,450
3,786
5,211
411,892
Something super weird is happening with Codex! My Luna reserve usage just crashed from 60-70% down to under 2% in two messages. Recent updates made it totally unusable. Time to switch back to Claude!? chatgpt.com/share/6ab59d89-4โ€ฆ
1
70
Finally the model card appeared. @wandb offers DeepSeek 4.1 flash with 1m context.
Hehe, so now they update the endpoint without updating the UI? @wandb released DeepSeek-V4.1-Flash a while ago! Don't know the pricing or Context, but I know its there (even though they didn't update the UI).
1
90
Steve ๐Ÿ‡บ๐Ÿ‡ธ retweeted
made these visuals with Opus 5.5 for people who get a headache comparing all the benchmarks. (between 6 Sol and 5.5 Opus) It seems like Opus 5.5 is the clear winner in terms of benchmarks, but 6 Sol is more Cost efficient.
2
7
1,334
I wanted sharing controls that get out of the way while you read. Share, copy link and email fold into a reading-progress ring as you scroll. Tap the percentage to bring them back. Inspired by Appleโ€™s new design.
1
64
Steve ๐Ÿ‡บ๐Ÿ‡ธ retweeted
Jev
166
1,289
17,898
697,157
Hehe, so now they update the endpoint without updating the UI? @wandb released DeepSeek-V4.1-Flash a while ago! Don't know the pricing or Context, but I know its there (even though they didn't update the UI).
I beat the @wandb intern to it again! Z.AI GLM 5.3 Flash is now live on serverless inference on WANDB. 1M Context, vision capable, 0.50$ per million output. Its also very fast. @CoreWeave
1
2
247
Only 128k context window? can anyone from @wandb confirm?
31
Sometimes your app needs a label, not a paragraph. JEV is TypeSafe's model for that. Send it text or JSON plus questions; get back choices, scores and probabilities your code can use. I'm not affiliated with TypeSafe or JEV.
1
2
274
Think of a support ticket: Choice: which team should handle it? Score: how urgent is it on your defined scale? Noul: does it ask for a refund? Noul returns 0โ€“1. A value of 0.91 means the model assigns a 91% probability to yes.
1
36
The useful part: one request can ask several independent questions about the same data, in parallel. Your code decides what happens next. Typed output isn't a guarantee of a correct answer. Docs: docs.typesafe.ai/prim itives
10
What happened to the @cursor_ai Origin? I haven't heard anything since launch day. It was never even offered to free users like they promised!
Origin, our code hosting platform, is now live. It's fast, easy to use, and deeply integrated with Cursor. Get started by syncing your repos from GitHub.
1
58
Grok 4.7 today or tomorrow?
1
70
My codex reset was automatically used? what the hell?
2
5
1,058
how the hell did I use my reset 10 minutes ago without knowing?
118
Was given early access to MiMo-X-Pro-Preview by @Xiaomi and its surprisingly a good model! Testing it some more and will update my finding in this thread.
1
203
I beat the @wandb intern to it again! Z.AI GLM 5.3 Flash is now live on serverless inference on WANDB. 1M Context, vision capable, 0.50$ per million output. Its also very fast. @CoreWeave
Hey, I beat the @wandb intern to it: DeepSeek V4 is now live on serverless inference. Letโ€™s see what they cooked.
2
1
4
1,357