Benchmark runs

Run explorer

Each row is one test: a real past change, and the files it needed. The squares show where each tool ranked those files: 20 squares for its top 20 picks, a filled square where it picked a needed file. More filled squares, further left, is better. Open a row to see both lists in full.

Tests
vs

    "Reported" holds the 30 tests per project shown on the findings page. "Tuning" holds the 10 per project used to make decisions. Times and costs are per test.