every run, one schema derived scores disclosed one number per surface

The measured board.

Every benchmark run behind this project — ours and the external rows we cite — rendered straight from the committed CSV (benchmarks/decision-model/results/board.csv). Filter it, sort it, and recompute the derived scores yourself; the formulas are on this page, the coefficients are yours to change, and nothing is averaged across benchmark surfaces unless you explicitly ask for it.

Live data

192 measured rows. No hand-typed numbers.

The table renders the same CSV the repo's docs of record are generated from — every row is a run with a committed artifact (see the footer for the source paths). Derived columns are marked derived and computed in your browser; everything else is the CSV verbatim.

The derived-score formulas — disclosed, adjustable, and not board numbers

Trust score derived
accuracy − λ · ECE, λ yours to set above (default 0.5). A model that is right but miscalibrated loses standing; a model that knows when it doesn't know keeps it. Rows without a measured ECE get no Trust score — never a guessed one.
Accuracy per GiB derived
accuracy ÷ model file size (GiB) — intelligence per gigabyte, the compression view. Only rows that publish a measured artifact size get one.
Rank & best-in-view
Ranked within each benchmark surface, never across them: a JevBench split score and a locked-suite score measure different things (the divergence is documented in docs/RESEARCH.md §14). "Best in view" above the table names one leader per surface, by your chosen metric.

These formulas are presentation, not measurement — the board numbers are the CSV. The only composite this project computes offline is the deployment composite (speed, resource, and accuracy-trust legs), which ships as its own rows in the CSV.

Benchmark runs with accuracy, calibration, latency, size, and derived scores. Click a column header to sort.