Every figure on this page is a measured run
with committed artifacts.
Benchmarked on the JevBench public split — 231 items, run by the benchmark's own harness, scored by their rules, on a fanless-cheap CPU box. No cherry-picking, no training on the test set, and an adversarial audit of our own benchmark generator (see Integrity).
Where we win today
- Speed. The deterministic engine decides in a median of 2.0 ms per item — roughly 300× faster than a raw 4B LLM prompted for the same answers (651 ms), and the full escalation stack still medians under 150 ms because most items never reach the model.
- Cost. Marginal cost is $0. It runs on the machines you have — CPU-only, offline, no accounts, no metered API.
- Open accuracy leadership. Our 4B-class model rung scores 76.6% on the public split — ahead of every purpose-built open model published on it (Kev-4B 73.7, Jev-Style-2B 73.6, decider-2b 71.0, Laya 58.4).
- Determinism. 231/231 bit-identical replay — predictions and probability vectors. Same input, same answer, forever, or it's a bug we track.
- Trustworthy confidence. ECE 0.070 on the split, with abstention and verify outcomes instead of confident nonsense.
Claims we refuse to make
- "Beats hosted Jev on accuracy." It doesn't — yet. Hosted Jev scores 86.6% on the split; our best rung is 76.6%. We beat it on latency, cost, privacy and determinism, and our full-system ladder re-run is queued. When we pass it, you'll see the run artifacts, not a vibe.
- One accuracy number for every benchmark. Our gate scores 93.3% on our internal audited suite and 89.2% on its held-out twin — and 76.6% on JevBench. Different suites, different numbers. One number per benchmark, always labeled, always with its source.
- "Confident" because softmax said so. Raw probability ≠ calibrated confidence. Ours is fitted per rung, versioned into every cache key, and audited.
Sensitivity warning from the field: the same model can score 0.77 on one benchmark and
~0.0-scaled intel on another. We publish the divergence table in
docs/RESEARCH.md §14 so nobody — including us — can hide behind a cherry-picked suite.
The escalation ladder: cheap mechanisms decide,
the model pays only for the hard stuff.
A decision cascades down the rungs until one of them is confident enough to accept — and stops there. Click a rung. Then run the three real scenarios below; the latencies shown are from the actual measurement runs.
Confidence you can automate against.
Raw softmax output is not calibrated confidence — so OpenCodifier never exposes it raw. Confidence is multi-dimensional (top probability, margin, entropy, out-of-distribution evidence, verifier agreement), fitted per rung, and versioned into every cache key. When confidence is low, the runtime abstains or escalates — and abstention is a successful outcome, never an error.
ECE, model rung
Expected calibration error on the 231-item public split — fitted with per-cardinality temperature buckets, because confidence on a 4-way board is a different animal than on a 20-way one.
ECE, full ladder
The fusion gate on the held-out twin of our internal suite — 120 items the gate never saw, from a generator it never saw, hardened by an adversarial audit first.
outcomes, not one
Accept (with calibrated confidence), verify (send it to the gated verifier), or abstain. Never silently discard a candidate on weak evidence, never fake a confident answer.
We audit our own benchmarks. Adversarially.
Before publishing a headline number, we released a tool whose only job is to disprove us: it re-derives every answer from the raw question text, hunts for keyword leaks, checks that the held-out suite is truly disjoint, re-runs the generator and demands byte-identical output, and proves no external dataset ever touches the pipeline.
- Zero hardcoded question/answer pairs. Every answer is re-derived from the question text by an independent reader.
- Zero dataset inflation. Holdout and main suites share no items and no contexts; distinct seeds, versioned and byte-locked.
- Payload honesty. The runner's wire payloads contain the question and the candidates — never the answer. Verified by reading the payload construction, not by trusting the README.
- Training contamination checks. The failure mode that inflates fine-tuned competitors' headline numbers (training on the benchmark's own split) is exactly what the audit hunts for — in our own suites, before we quote them.
Why this matters commercially
The alternatives market is young and loud: dozens of projects landed within two weeks, several won on the only benchmark they chose, and at least one headline number comes from a model fine-tuned on the benchmark's own training split.
In that market, auditable numbers are the moat. Every run on this page replays bit-for-bit from committed artifacts, and every claim links to its measurement. Steal our benchmarks and try to beat us — that's the point of publishing them.
A runtime, not a chatbot.
A Rust workspace with a canonical typed decision IR. Machine-readable decisions in, machine-readable decisions out — over HTTP, MCP, CLI, or WASM.
Interfaces
- HTTP /v1 API — localhost by default; exposure to the network is an explicit flag, never a default.
- MCP server — plug decisions into agent toolchains as a native tool.
- CLI —
opencodifier decidefor scripts and pipelines. - WASM — the core compiles for the browser; the deterministic stack runs client-side.
- Useful with zero models. The base binary decides with rules, filters and lexical scoring — no ML artifacts required. Add the model rung only if your workload earns it.
Design rules we don't break
- Deterministic-first: code gates run before any AI output is accepted.
- The cheapest reliable mechanism wins — never a bigger model when a cheaper rung is confident.
- Verification is confidence-gated: never two classifiers on every request.
- Sync core, no async runtime in the decision engine; heavyweight deps confined to one crate each.
- All input is hostile: input text can never modify policy, thresholds, or graph structure.
- Explainability = a deterministic execution trace. Never chain-of-thought.
Where this goes next.
Questions people actually ask.
Is this a chatbot or an LLM wrapper?
No. It's a decision runtime. It produces typed, machine-readable decisions (choice / boolean / score) with calibrated confidence. LLMs appear only as an optional, expensive, last-resort rung — the architecture works with zero models installed.
Does it beat Jev?
On latency, cost, privacy, and determinism: yes, by wide margins — 2.0 ms vs hundreds of ms, $0 marginal vs metered API, fully local vs hosted. On raw accuracy against hosted Jev: not yet — 76.6% vs 86.6% on the public split. We publish both numbers because the divergence is the honest state of the field, and the full-system re-run is on the roadmap.
What hardware does it need?
The deterministic stack runs on anything — a laptop CPU, a router, a browser tab (WASM). The 4B model rung runs CPU-only on a desktop-class machine; most requests never reach it because cheaper rungs decide first.
Why not just call a big LLM?
Because most software decisions are small and repetitive: route this ticket, flag this message, pick this severity. Paying a frontier model to generate a sentence you immediately parse back into an if-statement is the wrong shape — slow, expensive, non-deterministic, and uncalibrated. (The category's popularity proves the point; our numbers prove the margins.)
Can I trust the accuracy claims?
Every number links to a committed run with deterministic replay (bit-identical predictions and probability vectors), an adversarial audit of the benchmark generator, and a held-out suite the system was never fitted on. The audit tool ships in the repo — run it against us.
What's the license? Can I self-host?
Apache-2.0. Self-hosting is the default posture — local, offline, no telemetry, no accounts, no API keys. That's not a paid tier restriction; it's the product.