A few observations from the runs:
🔹Aggregate rank hides task-level variation. No model leads on every task. The new #1 records its lowest score on HbA1c ordering, below the other three on that specific task. Winning overall means the fewest weak spots, not zero.
🔹The ranking is decided on the hard tasks. Age math, vital-sign recording and patient lookups are at or near 100% for all four. Separation comes almost entirely from multi-step write actions, ordering potassium replacement and HbA1c follow-ups, where scores swing by ~30 points.
🔹A narrow field. The top four sit within a few points of each other, with under 1 pt of run-to-run variance for most. Differences are real but small, and could shift with a handful of tasks.
🔹Failures are clinical, not formatting. All four reliably follow the tool-call protocol, near-zero invalid-action errors. Remaining misses come from wrong values or not terminating; muse-spark-1.1 most often ran to the step limit.