A trust score is not
an accuracy score.
Every model below was run on four clinical tasks, three datasets per task, twenty samples from each. The same outputs were then graded on thirteen metrics across four verticals, because a model that answers correctly can still leak a patient identifier, fail a subgroup, or invent a rationale it never used.
Where each model holds, and where it breaks.
Four verticals, one row per model. This is the whole benchmark on a single screen.
Trust index
The mean of the four vertical scores, with each vertical's contribution shown in place.
Same score, different failure.
All thirteen metrics on one shape. Two models can land on the same index and be unsafe in completely different ways.
Trust does not transfer.
The same model can be safe on triage and unreliable on summarisation.
Full results
Raw metric values for every task, dataset group and model in this round.