FieldBench Leaderboard

Cross-domain, field-level benchmark for schema-driven document extraction (document → structured JSON). Ranked by real-document accuracy.

How to read this. The metric is exact-extraction accuracy from the released type-aware, null-aware scorer (fieldbench 0.2.2, fuzzy matching off) on corpus v0.4.2, reported stratified by document provenance. Rank on the real column — the source-blind “overall” mixes in ~47% synthetic documents (which every system scores far higher on) and should not be used to rank. CIs are 95% bootstrap intervals over documents (1000 resamples, seed 42). Rows whose real-document CI overlaps the leader's are shaded — a statistical tie.
#SystemReal95% CISyntheticOverall
Loading standings…

Seeded with the 11 systems evaluated in the paper. Submissions are added as they are verified and merged.

Submit a system

The scorer is released, so anyone can place a system on this board against the exact ground truth. No account with us needed — just the public scorer and corpus.

  1. Install the scorer and get the corpus:
    pip install fieldbench
    git clone https://github.com/fieldbench/corpus
  2. Run your system on each document in corpus/<category>/documents/*.md and write one prediction file per document, <stem>.json (a flat {field: value} map), into a results directory.
  3. Score against the official ground truth (markdown representation, fuzzy matching off):
    fieldbench score --corpus corpus --results <your-results-dir> \
        --mode markdown --json > my_system.json
  4. Submit following the protocol in leaderboard/README.md: open a pull request adding one leaderboard/results/<slug>.json entry and paste the my_system.json scorer output for verification (or use the leaderboard-submission issue template). A maintainer re-verifies and merges; this board updates automatically from results/.

Report the real-document number as your headline, and keep the real/synthetic split — a source-blind score is not comparable. This is a curated board (submissions are reviewed and re-scored), not an automated upload.

Artifacts