To physicians: never believe any study or post that confidently generalizes "specialized medical models" vs "general-purpose models" in clinical AI.
No matter who says it. No matter where it is published.
Before treating any result/headline as evidence or fact, ask a few basic questions:
1) What benchmarks or evaluation datasets were used in the study?
2) Which AI foundation models were tested?
3) Are the data and methodologies, including prompts and grading criteria, public and/or reproducible?
4) Was the comparison run with the same prompts, setup, grading method, and number of trials?
5) Who graded the outputs: clinicians, LLM judges, or vibes?
On this particular topic, general-purpose vs specialized medical models is still very much an open question. We need more rigorous, transparent, independent, and continuously updated validation before making broad conclusions beyond what has actually been tested.