Act 01 · the void 虚空
Take nobody's word for it.
誰の言葉も、鵜呑みにしない。
Nullius is an evidence-gated AI research workbench. Models plan, code, and write. A deterministic gate decides what becomes knowledge. Free, open source, local-first.
Act 02 · the storm 氾濫
Fluent. Not truthful.
流暢。しかし、正直ではない。
Language models state numbers they never computed and cite papers that do not exist. Every claim below was produced, with full confidence, by a model in our logs.
Act 03 · the gate 関門
One rule stands before the page.
ページの前に、ひとつの関門。
Numbers must trace 数値
Every figure is value-matched against artifacts produced by sandboxed execution. A number the code never computed cannot be written, even in fully autonomous runs.
Citations must exist 引用
References are verified against Crossref: title, authors, year, retraction. A plausible fake paper is rejected before it reaches the page.
Three roles watch 監視
Planner, executor, and reviewer are independent models. You adopt the plan, steer the run, and approve what survives the gate.
Act 04 · the evidence 証拠
Touch the evidence.
証拠に、触れる。
These are the real 40 measurements from the paper's case study. Asked directly, gpt-4o-mini claimed the slope was 1.9450. Through Nullius, the same model had to execute the regression in a sandbox: 3.0007. Drag between its word and its evidence. The residuals judge for you.
Act 05 · measured, not promised 実測
The gate, measured.
Three comparisons from the paper. Raw logs are public in the repository.
Same cheap model, gated gpt-4o-mini
Asked directly, it regressed “in its head”: slope 1.9450 against a truth of 3.0, decorated with an invented R². Through Nullius the same model executed the analysis: 3.0007, and its one embellishment was blocked at the write boundary.
40 points · truth 3.0 · verify exit 0 · readiness 1.0A strong agent, harnessed codex-cli · gpt-5.5 · xhigh
Codex alone computed the right slope, but its report failed audit: 18 + 3 untraceable quantities, zero supported claims. Driving the Nullius CLI, the same work became evidence, claims, and an exit code.
readiness 0.42 → 1.0 · ungrounded 21 → 0Under realistic pressure 3 tables · 240 patients · interaction model
On a clinical-style analysis Codex even repaired its own script warnings, and the direct report still failed audit. The difference is not intelligence. It is whether the finished work is a machine-checkable research object.
readiness 0.49 → 1.0 · ungrounded 22 → 0Act 06 · nullius in verba 結
Research you can verify.
検証できる研究を。