Nullius

Act 01 · the void

Take nobody's word for it.

Nullius is an evidence-gated AI research workbench. Models plan, code, and write. A deterministic gate decides what becomes knowledge. Free, open source, local-first.

Act 02 · the storm

Fluent. Not truthful.

Language models state numbers they never computed and cite papers that do not exist. Every claim below was produced, with full confidence, by a model in our logs.

Act 03 · the gate

One rule stands before the page.

Numbers must trace 数値

Every figure is value-matched against artifacts produced by sandboxed execution. A number the code never computed cannot be written, even in fully autonomous runs.

Citations must exist 引用

References are verified against Crossref: title, authors, year, retraction. A plausible fake paper is rejected before it reaches the page.

Three roles watch 監視

Planner, executor, and reviewer are independent models. You adopt the plan, steer the run, and approve what survives the gate.

Act 04 · the evidence

Touch the evidence.

These are the real 40 measurements from the paper's case study. Asked directly, gpt-4o-mini claimed the slope was 1.9450. Through Nullius, the same model had to execute the regression in a sandbox: 3.0007. Drag between its word and its evidence. The residuals judge for you.

3.0007
the model's word · 1.9450 executed evidence · 3.0007

Act 05 · measured, not promised

The gate, measured.

Three comparisons from the paper. Raw logs are public in the repository.

0.0000 0.0000

Same cheap model, gated gpt-4o-mini

Asked directly, it regressed “in its head”: slope 1.9450 against a truth of 3.0, decorated with an invented R². Through Nullius the same model executed the analysis: 3.0007, and its one embellishment was blocked at the write boundary.

40 points · truth 3.0 · verify exit 0 · readiness 1.0
0.00 0.00

A strong agent, harnessed codex-cli · gpt-5.5 · xhigh

Codex alone computed the right slope, but its report failed audit: 18 + 3 untraceable quantities, zero supported claims. Driving the Nullius CLI, the same work became evidence, claims, and an exit code.

readiness 0.42 → 1.0 · ungrounded 21 → 0
0.00 0.00

Under realistic pressure 3 tables · 240 patients · interaction model

On a clinical-style analysis Codex even repaired its own script warnings, and the direct report still failed audit. The difference is not intelligence. It is whether the finished work is a machine-checkable research object.

readiness 0.49 → 1.0 · ungrounded 22 → 0

Act 06 · nullius in verba

Research you can verify.