SAME QUESTION. DIFFERENT PATHS.
Every decision, in focus.
Explore how medical models answer, how certain they are, and how long they take.
OBSERVATORY 01
Each model works through the same queue at its own measured speed. The preview rate only controls the question/image viewer. Cards show each model’s latest completed case; their metrics may cover different numbers of questions. Saved-measurement replay, not live inference.
Green marks the reference answer for this case.
Three inference paths
Measured predictions · individual model progressDecision stream
LIVE · ONE ROW PER DECISIONEvery decision a model makes, as it completes cases — the answer it picked, its confidence, and the measured latency. Newest first.
Completed samples over time
SHARED CLOCK · INDEPENDENT WORKERSOne sequential worker per model, using saved per-case latency. No waiting for another model. Original hardware and batching conditions still apply.
The stream tells the story.
Cumulative metrics by scored sample. Hover to inspect a point.
Accuracy & confidence
Time to a decision
CUMULATIVE MEAN · MSRecorded timings retain their original hardware and batching conditions.
Stream scorecard
| Model | Scored | Accuracy ↑ | Confidence | Mean latency ↓ | P95 latency ↓ | Brier ↓ | ECE ↓ |
|---|
How to read this experiment
The 1,000-step SLAKE stream cycles through 355 unique held-out questions, alternating image groups. Repeats are labeled and are not independent measurements. MedQA uses 1,000 unique held-out questions.
Aligned comparison uses the same completed sample prefix for all models. Throughput race replays each model independently using cumulative saved latency, starting with zero completed samples on one shared clock. Pausing freezes that clock; resuming continues it. Preview speed and scrubbing do not change model throughput. The race estimates sequential throughput from original measurements; it is not a new hardware benchmark or live inference.
Race cards show different cases. Use “Inspect latest” to view the image, question, and options for a model’s most recent result. Quality plots use scored sample index; the throughput plot uses elapsed replay time.
Confidence is the predicted answer’s probability. Invalid AR outputs count as errors. Saved logits use their calibration temperature, and collected probabilities are preserved. Brier sums squared errors; ECE uses 10 equal-width confidence bins. All results come from stored measurements, including your collected MedJev run.
MedJev is text-only and appears in MedQA. SLAKE compares AR Gen, AR Head, AR LL, and NAR (Laya-Med). Scans are preloaded and decoded once before playback, then reused across questions.
How each path reaches a decision
The inference mechanism behind every model in this comparison — from the first prompt to the final choice.