Koragraph Benchmarks
How we measure against everyone else.
Every head to head we run, published with the corpus, the questions, the gold sets and the scorer, so the other system’s column can be reproduced by anyone who doubts it.
Runs
1 benchmarkWhat is next
Larger repositories, more systems, and the same rule every time: an independent referee rather than either side’s own definition of a correct answer, the corpus pinned to commit SHAs, and every input published so the comparison can be run against us. A rerun on a new corpus is published as a new benchmark rather than as an edit to an old one, so a number you saw once does not change underneath you.
