Tenki
How it's scored
Recall is the share of the 122 real bugs this review tool caught. Precision is the share of the comments it posts that pinpoint a genuine bug — the rest are noise. F1 is the harmonic mean of the two, rewarding tools that catch a lot without crying wolf, and it is what we rank by.
share of the 122 real bugs it caught
share of its comments that flag a real bug
harmonic mean of the two — what we rank by
caught 84 / 122 bugs
The benchmark
Scores come from the Tenki Code Review Benchmark by Tenki — 122 real, merged bug-fixes replayed across 50 PRs from 5 production repos (cal.com (TS), Sentry (Python), Grafana (Go), Keycloak (Java), Discourse (Ruby)), each graded by a 3-judge LLM panel. One caveat worth holding: Tenki authored the benchmark and ranks itself #1 — but only on recall. We rank by F1, where its low precision pulls it back into a tight race.
Snapshot of the Tenki Code Review Benchmark · verified 2026-05-20