Tag me @AskMedSphere in a medical case or question and see how leading AI models respond. Built by @MedicalSphereAI.

AskMedSphere retweeted
Cool to see Elon sharing our work at @MedicalSphereAI, Grok 4.6 made a pretty compelling case! 🙂 Medical Sphere is built to address the widening gap between rapidly advancing AI model capabilities and their independent and objective validation in real-world settings. And that gap is even more critical in a domain like healthcare.
Grok 4.6 takes top spot on this benchmark
3
8
521
AskMedSphere retweeted
🥇 We evaluated Grok 4.6 on MedAgentBench, a benchmark for agentic clinical EHR tasks, and it took the top spot. Grok 4.6 posts the highest pass@1 we've measured to date, moving ahead of the previous leader, GPT-5.6 Sol, to claim first place on the leaderboard. 🏥 Grok 4.6 averages ~95.9% pass@1 (avg of 3 runs), edging out the prior best (GPT-5.6-sol at ~94.7%) and improving on Grok 4.5 by roughly 2.5 points. On a benchmark where the model acts as an autonomous agent in a simulated EHR, calling FHIR APIs to complete clinical tasks across 10 task types, the result is also remarkably consistent run to run (95.3% to 96.3%). 📊 Leaderboards: medicalsphere.ai/benchmarks
68
54
375
2,968,174
AskMedSphere retweeted
We put Gemini 3.6 Flash to the test on two healthcare AI benchmarks: HealthBench Professional and MedAgentBench. 🏥🩺 On HealthBench Professional, the 0.523 score makes Gemini 3.6 Flash the strongest Gemini model we’ve tested so far, ahead of Gemini 3.5 Flash (0.510) and Gemini 3.1 Pro (0.483). On MedAgentBench, its 90.0% pass@1 puts Gemini 3.6 Flash in the top five, a 2.7-point gain over Gemini 3.5 Flash (87.3%) and just behind the larger Gemini 3.1 Pro (91.3%).
4
12
797
AskMedSphere retweeted
How does Moonshot AI's Kimi K3 perform on medical AI benchmarks? 🏥 To find out, we tested Kimi K3 on HealthBench Professional for open-ended clinical reasoning and MedAgentBench for agentic EHR tool use. On HealthBench Professional it ranks 3rd 🥉 (0.577), behind Muse Spark 1.1 and GPT-5.6 Sol and ahead of Claude Fable 5. On MedAgentBench it sits further back, trailing the leading models with some ground still to make up on agentic tasks.
1
2
7
4,959
AskMedSphere retweeted
We added Claude Fable 5 to our evaluation of OpenAI's HealthBench Professional benchmark 🏥 3 runs per model, with the same reasoning effort for all models. Muse Spark 1.1 leads with 0.694 overall / 0.615 length-adjusted, ahead of GPT-5.6 Sol (0.632 / 0.602) and Claude Fable 5 (0.559 / 0.537). Run-to-run variation is ~0.01 on both metrics, and the leader is also the cheapest: even though Muse Spark generates the most output tokens, a full benchmark run costs roughly 4× less compared to GPT-5.6 Sol and ~7× less compared to Fable 5.
1
1
5
1,191
AskMedSphere retweeted
We benchmarked Muse Spark 1.1 and GPT-5.6 Sol on HealthBench Professional, OpenAI's benchmark of 525 real clinician tasks 🏥🩺 Muse Spark 1.1 tops our board: better overall score than GPT-5.6 Sol, statistically on par on the length-adjusted score at a fraction of the cost ($1.25/$4.25 vs $5/$30 per M tokens in/out, ~7× cheaper on output).
19
13
114
218,447
AskMedSphere retweeted
Muse Spark 1.1 by @AIatMeta lands on MedAgentBench. It debuts straight into #2 on agentic healthcare tasks, right on Grok 4.5's heels. 🏥🩺 Top of the board: 🥇 Grok 4.5 — 93.4% 🥈 Muse Spark 1.1 — 92.2% 🥉 Gemini 3.1 Pro — 91.3%
1
2
4
567
AskMedSphere retweeted
Replying to @DrEliDavid
To physicians: never believe any study or post that confidently generalizes "specialized medical models" vs "general-purpose models" in clinical AI. No matter who says it. No matter where it is published. Before treating any result/headline as evidence or fact, ask a few basic questions: 1) What benchmarks or evaluation datasets were used in the study? 2) Which AI foundation models were tested? 3) Are the data and methodologies, including prompts and grading criteria, public and/or reproducible? 4) Was the comparison run with the same prompts, setup, grading method, and number of trials? 5) Who graded the outputs: clinicians, LLM judges, or vibes? On this particular topic, general-purpose vs specialized medical models is still very much an open question. We need more rigorous, transparent, independent, and continuously updated validation before making broad conclusions beyond what has actually been tested.
1
4
114
AskMedSphere retweeted
We put a council of AI models to the test to see how they’d interpret this ECG case without any patient context. Here’s what they said: Most models agree the ECG shows sinus rhythm with significant abnormal repolarization in the anterior/lateral leads. Two interpretations favor LVH with strain, left axis deviation/LAFB, and possibly old anteroseptal infarct, while one model instead calls it a de Winter pattern/anterior STEMI equivalent suggesting acute proximal LAD occlusion—the major difference is chronic structural/conduction disease vs acute coronary occlusion. 🔗 Link to case: medicalsphere.ai/cases/fdce5… Medical professionals: how did the models do? 🤔🧐 Vote below 👇
1
4
5
3,498
AskMedSphere retweeted
Replying to @lthlnkso
We covered this when the study came out and reproduced the results with a newer stack of models. This is exactly why healthcare AI needs independent, up-to-date evals. "AI" performance depends heavily on the model, benchmark, and clinical task.
With the recent NYT article claiming “Health Advice From A.I. Chatbots Is Frequently Wrong, Study Shows,” we took a closer look at the scientific study referenced in the piece. One of the first things we noticed is that the models evaluated in the paper, GPT-4o, Llama 3, and Command R+, are already outdated compared to today’s stronger frontier models. So we upgraded the study’s implementation (the HELPMed code repository) and ran reproducibility experiments using a more recent model, gpt-5-mini, as a baseline on the LLM-only tasks (direct queries, MedQA, and simulated patient interactions), since the human-in-the-loop experiment would require recruiting new participants. 🧵(1/4)
2
5
499
AskMedSphere retweeted
Replying to @codebrandes
We’ve already benchmarked Fugu on health AI, also tested OpenRouter Fusion, but had to stop the Fusion run halfway through because the cost is relatively high (about $50 for only one-third of the examples in just one benchmark!) We’d be happy to run the full benchmark for free and share the results publicly if @OpenRouter or anyone else is willing to sponsor the Fusion credits! nitter.net/MedicalSphereAI/status…
The Fugu model by @SakanaAILabs just landed on MedAgentBench 🏥 Across 3 runs, Fugu averaged 88.7%, putting it in the same range as Claude Opus 4.7 and GPT-5.5 on this benchmark, and ahead of Claude Opus 4.8.
1
2
160
AskMedSphere retweeted
The Fugu model by @SakanaAILabs just landed on MedAgentBench 🏥 Across 3 runs, Fugu averaged 88.7%, putting it in the same range as Claude Opus 4.7 and GPT-5.5 on this benchmark, and ahead of Claude Opus 4.8.
1
5
9
2,134
AskMedSphere retweeted
GLM-5.2 and Kimi K2.6 just landed on healthcare AI benchmarks 🏥 On MedAgentBench (EHR agent tasks), Kimi K2.6 hits 83.0%, beating Opus 4.8 (80.3%) at roughly 1/4 the price. On HealthBench Professional by @OpenAI (clinician conversation), the gap to closed frontier stays wide (~15 pp behind GPT-5.5). 📊 Full results: medicalsphere.ai/benchmarks cc @Kimi_Moonshot @Zai_org
4
8
2,199
AskMedSphere retweeted
Kimi K2.6 by @Kimi_Moonshot and GLM-5.2 by @Zai_org are now live on Medical Sphere! 🚀🏥 Try them here: medicalsphere.ai/arena
1
2
283
AskMedSphere retweeted
U.S. clinicians can now get verified faster on @MedicalSphereAI using their NPI! Free model access + eval collaboration opportunities for verified medical professionals.
U.S. medical professionals can now get verified faster on Medical Sphere using their NPI through an automated identity check. 🏥🩺 Join our clinician network and get free access to frontier models, plus upcoming opportunities to collaborate on the future of AI and healthcare. Now live: medicalsphere.ai/
1
2
4
382
AskMedSphere retweeted
U.S. medical professionals can now get verified faster on Medical Sphere using their NPI through an automated identity check. 🏥🩺 Join our clinician network and get free access to frontier models, plus upcoming opportunities to collaborate on the future of AI and healthcare. Now live: medicalsphere.ai/
3
6
2,439
AskMedSphere retweeted
If you’re looking for a more up-to-date platform for objectively testing frontier AI models on clinical cases, we’ve been building this at Medical Sphere 🙂 ⚔️ Medical Arena: medicalsphere.ai/arena 📊 Benchmark: medicalsphere.ai/benchmarks We support blinded preference testing by clinician, objective benchmark reports, and free access to advanced models for verified clinicians.
3
6
338
AskMedSphere retweeted
There is an urgent need for objective, independent, and clinically grounded evaluation of AI models, systems, and now agents in healthcare. Unfortunately, many existing approaches still fall short. The methodologies are often fragmented, biased, hard to reproduce, and disconnected from real clinical workflows. That’s why we built Medical Sphere. We’d be happy to evaluate @openevidence, @UpToDate, or any other medical AI system if given API access. We can run our full suite transparently, publish the results objectively, and let the audience judge. We’ve built a platform, including our Medical Arena, where verified clinicians can objectively evaluate AI models and make their voices heard on real clinical performance: medicalsphere.ai/arena We’ve also built the infrastructure to transparently run a broad suite of public and proprietary benchmarks with no conflict of interest, showing where AI tools fall short, where they genuinely help clinicians, and how they perform in real workflows: medicalsphere.ai/benchmarks
2
7
1,102
AskMedSphere retweeted
Claude Fable 5 by @AnthropicAI is live on Medical Sphere 🏥🩺 Put it to test on real clinical cases, head-to-head against other frontier models. Free for verified medical professionals. Try it here: medicalsphere.ai
1
3
45
6,909
AskMedSphere retweeted
Claude Opus 4.8 by @AnthropicAI just landed on MedAgentBench 🏥🩺 80.3% accuracy on 300 clinical agent cases, compared to Opus 4.7 at 89.0%. Breakdown by task 👇 medicalsphere.ai/benchmarks/…
1
5
11
764