Apache-2.0 local-first no API keys CPU-only

Decisions,
not paragraphs.

OpenCodifier turns unstructured state into typed, machine-actionable decisions — choice, boolean, score — with calibrated confidence and honest abstention, on hardware you already own. A large language model is the expensive last resort here, not the default.

See the measured numbers How the ladder works
0 ms
engine p50 per decision
JevBench public split, CPU-only
0%
best published open decision model
on the 231-item public split
$0
marginal cost per decision
local, offline, self-hosted
0/231
deterministic replay — every item,
every time, bit-for-bit
0
calibration error (ECE)
of our model rung on the split
The numbers

Every figure on this page is a measured run
with committed artifacts.

Benchmarked on the JevBench public split — 231 items, run by the benchmark's own harness, scored by their rules, on a fanless-cheap CPU box. No cherry-picking, no training on the test set, and an adversarial audit of our own benchmark generator (see Integrity).

Where we win today

  • Speed. The deterministic engine decides in a median of 2.0 ms per item — roughly 300× faster than a raw 4B LLM prompted for the same answers (651 ms), and the full escalation stack still medians under 150 ms because most items never reach the model.
  • Cost. Marginal cost is $0. It runs on the machines you have — CPU-only, offline, no accounts, no metered API.
  • Open accuracy leadership. Our 4B-class model rung scores 76.6% on the public split — ahead of every purpose-built open model published on it (Kev-4B 73.7, Jev-Style-2B 73.6, decider-2b 71.0, Laya 58.4).
  • Determinism. 231/231 bit-identical replay — predictions and probability vectors. Same input, same answer, forever, or it's a bug we track.
  • Trustworthy confidence. ECE 0.070 on the split, with abstention and verify outcomes instead of confident nonsense.

Claims we refuse to make

  • "Beats hosted Jev on accuracy." It doesn't — yet. Hosted Jev scores 86.6% on the split; our best rung is 76.6%. We beat it on latency, cost, privacy and determinism, and our full-system ladder re-run is queued. When we pass it, you'll see the run artifacts, not a vibe.
  • One accuracy number for every benchmark. Our gate scores 93.3% on our internal audited suite and 89.2% on its held-out twin — and 76.6% on JevBench. Different suites, different numbers. One number per benchmark, always labeled, always with its source.
  • "Confident" because softmax said so. Raw probability ≠ calibrated confidence. Ours is fitted per rung, versioned into every cache key, and audited.

Sensitivity warning from the field: the same model can score 0.77 on one benchmark and ~0.0-scaled intel on another. We publish the divergence table in docs/RESEARCH.md §14 so nobody — including us — can hide behind a cherry-picked suite.

How it works

The escalation ladder: cheap mechanisms decide,
the model pays only for the hard stuff.

A decision cascades down the rungs until one of them is confident enough to accept — and stops there. Click a rung. Then run the three real scenarios below; the latencies shown are from the actual measurement runs.

Pick a scenario to watch a request ride the ladder.
Calibration

Confidence you can automate against.

Raw softmax output is not calibrated confidence — so OpenCodifier never exposes it raw. Confidence is multi-dimensional (top probability, margin, entropy, out-of-distribution evidence, verifier agreement), fitted per rung, and versioned into every cache key. When confidence is low, the runtime abstains or escalates — and abstention is a successful outcome, never an error.

0.070

ECE, model rung

Expected calibration error on the 231-item public split — fitted with per-cardinality temperature buckets, because confidence on a 4-way board is a different animal than on a 20-way one.

0.103

ECE, full ladder

The fusion gate on the held-out twin of our internal suite — 120 items the gate never saw, from a generator it never saw, hardened by an adversarial audit first.

3

outcomes, not one

Accept (with calibrated confidence), verify (send it to the gated verifier), or abstain. Never silently discard a candidate on weak evidence, never fake a confident answer.

Integrity

We audit our own benchmarks. Adversarially.

Before publishing a headline number, we released a tool whose only job is to disprove us: it re-derives every answer from the raw question text, hunts for keyword leaks, checks that the held-out suite is truly disjoint, re-runs the generator and demands byte-identical output, and proves no external dataset ever touches the pipeline.

  • Zero hardcoded question/answer pairs. Every answer is re-derived from the question text by an independent reader.
  • Zero dataset inflation. Holdout and main suites share no items and no contexts; distinct seeds, versioned and byte-locked.
  • Payload honesty. The runner's wire payloads contain the question and the candidates — never the answer. Verified by reading the payload construction, not by trusting the README.
  • Training contamination checks. The failure mode that inflates fine-tuned competitors' headline numbers (training on the benchmark's own split) is exactly what the audit hunts for — in our own suites, before we quote them.

Why this matters commercially

The alternatives market is young and loud: dozens of projects landed within two weeks, several won on the only benchmark they chose, and at least one headline number comes from a model fine-tuned on the benchmark's own training split.

In that market, auditable numbers are the moat. Every run on this page replays bit-for-bit from committed artifacts, and every claim links to its measurement. Steal our benchmarks and try to beat us — that's the point of publishing them.

Build with it

A runtime, not a chatbot.

A Rust workspace with a canonical typed decision IR. Machine-readable decisions in, machine-readable decisions out — over HTTP, MCP, CLI, or WASM.

Interfaces

  • HTTP /v1 API — localhost by default; exposure to the network is an explicit flag, never a default.
  • MCP server — plug decisions into agent toolchains as a native tool.
  • CLI — opencodifier decide for scripts and pipelines.
  • WASM — the core compiles for the browser; the deterministic stack runs client-side.
  • Useful with zero models. The base binary decides with rules, filters and lexical scoring — no ML artifacts required. Add the model rung only if your workload earns it.

Design rules we don't break

  • Deterministic-first: code gates run before any AI output is accepted.
  • The cheapest reliable mechanism wins — never a bigger model when a cheaper rung is confident.
  • Verification is confidence-gated: never two classifiers on every request.
  • Sync core, no async runtime in the decision engine; heavyweight deps confined to one crate each.
  • All input is hostile: input text can never modify policy, thresholds, or graph structure.
  • Explainability = a deterministic execution trace. Never chain-of-thought.
Roadmap

Where this goes next.

Adversarial benchmark audit — shipped. Independent re-derivation of every suite answer; hardened holdout re-measured.
Held-out validation — shipped. Full-system ladder: 0.892 out-of-sample with zero gate drift, +9.2 pp over the previous gate on identical hardware.
Full-system JevBench re-run — the fusion gate on the official public split, run by the benchmark's own harness.
Official third-party board submission — 1,500-decision self-hosted run, artifacts published, independently checkable.
Smaller, faster model rung — distilling our findings into a compact model targeting the same accuracy at a fraction of the compute.
Sub-100 ms ONNX encoder rungs — a middle rung between lexical and the 4B model, CPU-only.
Release engineering — binaries for Linux, macOS, Windows, ARM and WASM, with release attestation like everything else we ship.
FAQ

Questions people actually ask.

Is this a chatbot or an LLM wrapper?

No. It's a decision runtime. It produces typed, machine-readable decisions (choice / boolean / score) with calibrated confidence. LLMs appear only as an optional, expensive, last-resort rung — the architecture works with zero models installed.

Does it beat Jev?

On latency, cost, privacy, and determinism: yes, by wide margins — 2.0 ms vs hundreds of ms, $0 marginal vs metered API, fully local vs hosted. On raw accuracy against hosted Jev: not yet — 76.6% vs 86.6% on the public split. We publish both numbers because the divergence is the honest state of the field, and the full-system re-run is on the roadmap.

What hardware does it need?

The deterministic stack runs on anything — a laptop CPU, a router, a browser tab (WASM). The 4B model rung runs CPU-only on a desktop-class machine; most requests never reach it because cheaper rungs decide first.

Why not just call a big LLM?

Because most software decisions are small and repetitive: route this ticket, flag this message, pick this severity. Paying a frontier model to generate a sentence you immediately parse back into an if-statement is the wrong shape — slow, expensive, non-deterministic, and uncalibrated. (The category's popularity proves the point; our numbers prove the margins.)

Can I trust the accuracy claims?

Every number links to a committed run with deterministic replay (bit-identical predictions and probability vectors), an adversarial audit of the benchmark generator, and a held-out suite the system was never fitted on. The audit tool ships in the repo — run it against us.

What's the license? Can I self-host?

Apache-2.0. Self-hosting is the default posture — local, offline, no telemetry, no accounts, no API keys. That's not a paid tier restriction; it's the product.