Groundtruth Dynamic Geological Benchmark
A fixed benchmark tells you what a model can do — not what it just learned from your corpus.
A benchmark harness that measures what a fine-tuned model actually learned from a specific corpus, not what it can do in general. Rather than a fixed question set, it ships a pipeline that reads an authenticated source corpus, extracts testable concepts, and writes questions with source-cited grading rubrics — then has an independent LLM judge score answers pointwise or head-to-head, with position-bias correction and contamination auditing. The v1.0.0 release includes a runnable geology edition built on 34 openly licensed mineral-deposit records from the Yudnamutana Copper district, South Australia, plus the corpus-to-rubric authoring pipeline itself, so the same method transfers to any domain.