SYSONE-MED / biomedia

System One · direct medical decisions

Same question. Five ways to decide.

How medical vision-language models reach a multiple-choice answer — five decision paths, from slow autoregressive writing to a single non-autoregressive scoring pass. Don't just read about it: play with each one and watch every decision it makes.

5 decision variants 7 medical datasets 5 interactive playgrounds saved measurements · no live GPU

Playground

Five ways to feel a decision model

Each one replays the same five variants — but from a different angle. Every action shows the decision it made, how confident it was, and how long it took.

—
Fastest decision (mean latency)
—
Slowest decision (mean latency)
—
Speed spread across variants
—
Best held-out accuracy (SLAKE)

The comparison

Five decision paths, one question

Each path's measured mean decision latency and held-out accuracy, ranked by speed.

From the paper

The evidence

Figures from the SysOne study.

Four decision formulations

Four decision formulations

AR generation, label likelihood, and learned heads on AR and NAR backbones — same inputs, different decision interfaces.

NAR direct decision architecture

Direct decision architecture

Text and image are encoded once, fused, and scored by a shared head — batched candidate scoring, no autoregressive decoding.

Accuracy versus latency

Accuracy vs. latency

Held-out accuracy against p50 decision latency across image and text tasks. Upper-left is better.

Calibration versus latency

Calibration vs. latency

ECE and NLL against latency. Lower-left is better — direct scoring stays calibrated without the generation cost.

Post-hoc test-time adaptation

Post-hoc test-time adaptation

Routing and conformal TTA over candidate scores — no weight updates, with a marginal coverage guarantee.

Visual dependence ablation

Visual dependence

Accuracy change when the image is missing or shuffled — how much each path actually relies on the visual input.

Under the hood

How each path reaches a decision

Autoregressive

AR Gen · AR LL

Writes text — up to 32 tokens, or exactly one — then parses it into a choice.

Learned head

AR Head

A frozen backbone feeds a small head that scores every option at once.

Non-autoregressive

NAR

Text and image are encoded, fused, and scored in a single forward pass.