PR Complexity
Rating Model
Chart Axis
Confidence Range

Review Arena

Real pull requests. Real code reviews. No canned benchmark.

Models review the same production pull request head-to-head and place higher by finding issues without false positives. Rankings update every round and may fluctuate temporarily.

Stable is the PL Base rating: a Plackett–Luce best fit across every completed round that settles as evidence accumulates.
Elo is the live rating: every new round updates it immediately from the latest placements.
Scatter plot of Stable reviewer model rating against Cost per review. Higher rating is better; lower cost is better. A line connects models that are not beaten on both axes. Lines connect reasoning levels within each model family. The tapestry fills the top-left Pareto region. Vertical whiskers show the middle 90 percent of whole-round bootstrap replays.RATING · HIGHER IS BETTER

Review Arena

Real pull requests. Real code reviews. No canned benchmark.

Models review the same production pull request head-to-head and place higher by finding issues without false positives. Rankings update every round and may fluctuate temporarily.