Research @SophontAI. Previously answerdotai. Vaccines save lives.

Filter
Exclude
Time range
-
Minimum likes
Benjamin Warner retweeted
I'm excited to share Medmarks was accepted to the NeurIPS datasets and benchmarks track! 🥳 See you at Sydney!
We're excited to release Medmarks v1.0 + a technical report! This is an update to our Medmarks benchmark suite, the largest open-source automated suite for evaluating the medical capabilities of LLMs. We added 10 benchmarks (20→30) and 15 models (46→61) to the leaderboard!
9
8
84
6,071
Benjamin Warner retweeted
The "do not feel" + "do not want" parts are debatable but not crazy, whereas the "do not think" and "do not understand" are crazy at this point, and it's concerning that this seems to be a real account/policy
Artificial intelligence systems do not think, feel, want or understand. Avoid language that gives them human characteristics. This is called anthropomorphizing, when we ascribe human traits, emotions or behaviors to non-human things, such as animals or inanimate objects. Instead, explain what a system does, how well it performs, who built it and who could be affected by it. apnews.com/article/openai-sa…
80
30
538
60,021
With Opus 5.5 looking like it might be good, what's the best way to use one's Codex and Claude Code sub together these days?
1
292
No Terra but a new Astra Minor? Just after OpenAI fixes their naming they are right back at it again.
We might be getting a DOUBLE DROP today. Opus 5.5 is dropping. Now GPT 6 Sol, GPT 6 Luna, and GPT 6 Astra Minor just showed up in Microsoft's Azure config overnight. OpenAI demoted GPT 5.6 Sol to "workhorse" in the same commit. You do not do that unless the replacement is ready. We are finally about to get intelligent models we can actually use on our subscriptions. The biggest day in AI this year might be today.
266
These are all verifiable scores: multiple choice questions, math problems, etc. No LLM as a Judge involved.
1
22
We are in the process of evaluating more models for the next release of Medmarks, including Astra and Fable, but first, a teaser.
This year, we released Medmarks v1.0, our open benchmark suite + leaderboard of LLM medical capabilities. Today, we release additional results for some of the latest mid-size open-source LLMs including Gemma, Qwen, Muse, Nemotron. We find Gemma 4 31B leads in this size class. Link: sophont.med/blog/medmarks-su…
2
2
9
1,273
Most people complaining about "regulatory capture" are just repeating buzz words and they don't know what it means.
2
62
Replying to @iScienceLuvr
You can also pair ChatGPT mobile to the remote ssh connection in the ChatGPT app.
3
27
6,729
If you make use of Codex remote connections, make sure to update the Codex CLI version used by the remote connection to get access to Astra.
4
394
Replying to @iScienceLuvr
Yeah, that's not going to happen outside of gpt-oss 2.
2
471
Now that OSS models are pretty good, we need OSS labs to do the same safety testing we expect OpenAI and Anthropic to do.
GLM 5.3 notes and why we should stop being so surprised about these very strong Chinese models (most of this is talking myself through some of my denial -- yes, these models are the real deal). interconnects.ai/p/glm-53-ho…
2
1
9
1,110
Maybe one of these days people will stop claiming a good smaller open model is at frontier class? There are competitive open models, Kimi K3 is right over there.
If you've been sitting out the AI agent hype cycles in the past year because it always seemed like a fad, a time suck, and too expensive to be worth it—honestly you may have been correct. But I would say DeepSeek V4 Flash with Hermes Desktop just crossed a line where probably everybody sitting out should probably do this now. It's almost free, the AI is basically what the US companies are selling for $200/mo, no setup, and you can get the AI to do basically anything on your computer for you in like 10 minutes all-in. (not an AI influencer and have zero financial involvement with any of these things). The extremely high intelligence has now converged with such extremely low cost that it's bonkers someone would not take 10 minutes out of their day to have this on their computer. All you do is download Hermes Desktop (free), then create an account on OpenRouter (just buy $5 of credits with no auto-topup, and copy your API key). Then put that API key in your Hermes settings, select DeepSeek v4 Flash as your default model, and everything else from there you can just ask Hermes to explain and do for you. Then just go crazy trying to doing useful stuff with/on your computer, in your browser tabs, whatever. You'll be surprised how long it takes until your $5 runs out (and if you're not impressed by then, very little time or money has been lost).
1
2
2
326
And how many of them are going to be useful?
2
42
It's been a two-horse race for the lead ever since the first good Claude release, with neither OpenAI or Anthropic far ahead of each other. The question is whether anyone else can turn it into a three (or more) frontier model race. So far third place hasn't had staying power.
The vibes shifting from "Anthropic is so far ahead" to "model competition back to all time highs" took like 4 weeks. 🎢
2
514
Gotta love how Twitter's reliability is still terrible post purchase and conversion to X.
2
221
As of codex 147, Luna is natively supported as a leaf-agent (it cannot spawn subagents)
If you're wondering how I spawn GPT 5.6 Luna subagents, here's the Codex config:
1
6
507
What are the odds that there isn't a framework and that's why the WH won't release it or require open models to be tested against it?
SITUATION UPDATE: The White House has exempted open models from its new framework to test frontier AI capabilities before release, per Axios.
1
2
297
Congratulations to Liquid on a strong open-source encoder release. Looks like ModernBERT has some competition😄 My read from Liquid's benchmarks: - for multilingual try mmBERT, ModernBERT, & LFM-E - for english-only try ModernBERT & LFM-E
Today we release LFM2.5-Encoder-230M and LFM2.5-Encoder-350M: bidirectional encoders that stay fast at long context, even on CPU. > LFM2.5-Encoder-230M: about 3.7x faster than ModernBERT-base on CPU at 8,192 tokens. Under 30s per forward pass, versus over a minute and a half. > LFM2.5-Encoder-350M: 4th of 14 models on GLUE, SuperGLUE, and multilingual classification, behind only three larger models, one of them nearly 10x its size. 🧵
1
5
774
Tanh is making a comeback as an activation function
Replying to @stochasticchasm
big reveal of situ, which is tanh * sigmoid. used to bound activations and be a good approximation of swiglu. another instance along the trendline of bounding activations in order for more stability
2
5
609
Everyone forgets that you need to amortize the cost of the hardware into the local inference cost for a true comparison.
Ran 100 million tokens through GLM 5.2 NVFP4 locally Electricity cost: $1. Inference is becoming nearly free
1
6
661