web4-bench

web4-bench is every situation, question, label and invariant from web4's conformance suites, published as a dataset with a scorer. It measures an engine exactly the way web4 does: it replays the engine's answers through the real planner and checks the page a visitor would get.

The challenge

Make a small open System One engine plan pages as well as Jev. An entry meets the challenge when both of these hold:

  • Small and open: open weights, at most 2 GB of them, and it answers offline on a laptop CPU.
  • As good as Jev on the held-out test split: its decision accuracy and its invariant pass rate are each at least those of the best Jev entry.

Distillation is one route: ask Jev the dev-split questions with your own key, then train a small model on its answers. The dataset itself ships labels only, with no Jev answers.

Leaderboard

SplitSitesItemsLabels
devCasa Lumbre (restaurant), Meridian Supply (database explorer)2771,277
testCasa Ribeira (hotel starter), Welcome starter132602

Dataset v1, content hash cd35b17ca8b0.

EngineTest decisionTest rawTest invariantsDev decisionOpen, smallAnswersChallenge
rules100.0%100.0%100.0%80.3%nopublished, re-scored in CI–
jev-1.13.099.7%100.0%100.0%96.2%noreported only, 2026-10-01–
  • Decision accuracy scores labels against the final page: after confidence gating, defaults, invariants and the rules fallback. It is what a visitor actually gets.
  • Raw accuracy scores the engine's answers before any gating.
  • Invariants are page properties every plan must keep, such as "never offer rooms to a guest who is already staying".

The rules engine replays each site's hand-written heuristics. Those cover the starter sites' fixtures completely, which is why it scores 100% on the test split. It is the floor web4 falls back to, not a System One engine, so it can't meet the challenge. The heuristic ablation measures what happens without those heuristics.

Jev's entry is reported only: its answers aren't redistributed, so CI can't re-score it. It was measured with Jev 1.13.0, 3 repeats, from the ablation's recordings.

The dataset

Dev splitThe two example sites: Casa Lumbre (a restaurant) and Meridian Supply (a database explorer).
Test splitThe two starters: Casa Ribeira (a hotel) and the welcome page. They are held out by site, so an engine is measured on sites it wasn't tuned on.
FormatOne item per situation: the state and questions in the System One wire format, ready to send, plus labels and invariants.
RulesThe test labels are public. Training on them, or tuning against test scores, disqualifies an entry.

Run it

The dataset lives in bench/data/v1 in the repository. Everything below runs from a clone:

git clone https://github.com/shd8/web4kit && cd web4kit && pnpm install
pnpm bench run --engine laya                      # or rules, jev (estimates spend first), local
pnpm bench score /tmp/web4-bench-*.jsonl.gz

--engine local sends the questions to any System One endpoint on your machine (W4_LOCAL_ENGINE_URL). The bench README has the item and submission formats, the scoring method, and how to add an entry with a pull request.

Try it on your site

pnpm create web4kit

Without a key, pages are planned by the offline rules engine. For a System One model, get a Jev key from the TypeSafe console and put it in .env.