SYSONE-MED / biomedia●   QA + VQA BENCHMARK

SAME QUESTION. DIFFERENT PATHS.

Every decision, in focus.

Explore how medical models answer, how certain they are, and how long they take.

INFERENCE
OBSERVATORY 01
EVALUATION SETSLAKEVisual question answering
UNIQUE QUESTIONS355Held-out test cases
MODELS COMPARED03Same question & choices
STREAM PROGRESS0%First pass
Throughput race
Preparing comparison…
THE INPUTSLAKE / TEST
MEDICAL IMAGEPreparing scans…
VISUAL QUESTION ANSWERING

Green marks the reference answer for this case.

Three inference paths

Measured predictions · individual model progress

Decision stream

LIVE · ONE ROW PER DECISION

Every decision a model makes, as it completes cases — the answer it picked, its confidence, and the measured latency. Newest first.

The stream tells the story.

Cumulative metrics by scored sample. Hover to inspect a point.

Accuracy & confidence

Time to a decision

CUMULATIVE MEAN · MS

Recorded timings retain their original hardware and batching conditions.

Stream scorecard

Model Scored Accuracy ↑ Confidence Mean latency ↓ P95 latency ↓ Brier ↓ ECE ↓
How to read this experiment

The 1,000-step SLAKE stream cycles through 355 unique held-out questions, alternating image groups. Repeats are labeled and are not independent measurements. MedQA uses 1,000 unique held-out questions.

Aligned comparison uses the same completed sample prefix for all models. Throughput race replays each model independently using cumulative saved latency, starting with zero completed samples on one shared clock. Pausing freezes that clock; resuming continues it. Preview speed and scrubbing do not change model throughput. The race estimates sequential throughput from original measurements; it is not a new hardware benchmark or live inference.

Race cards show different cases. Use “Inspect latest” to view the image, question, and options for a model’s most recent result. Quality plots use scored sample index; the throughput plot uses elapsed replay time.

Confidence is the predicted answer’s probability. Invalid AR outputs count as errors. Saved logits use their calibration temperature, and collected probabilities are preserved. Brier sums squared errors; ECE uses 10 equal-width confidence bins. All results come from stored measurements, including your collected MedJev run.

MedJev is text-only and appears in MedQA. SLAKE compares AR Gen, AR Head, AR LL, and NAR (Laya-Med). Scans are preloaded and decoded once before playback, then reused across questions.

UNDER THE HOOD

How each path reaches a decision

The inference mechanism behind every model in this comparison — from the first prompt to the final choice.