Deterministic
The effectiveness computation: which comments exist, which were followed by a code change in the PR, and the aggregations by reviewer, team, and period. Running it twice over the same history produces the same number.
// METHODOLOGY
A measurement only deserves trust if it can be questioned. This page defines Releezy Guardian’s central metric, separates deterministic computation from model assistance, and lists the limits we know about.
A review comment counts as effective when it is followed by a real code change in the same pull request, between the comment and the merge. It is a comment-to-change conversion measure, computed from Git history, with the same rule for every reviewer: human, AI tool, or Releezy’s own agents.
The name matters: we measure whether the comment was followed by a change, not whether the comment was right. The two correlate, but they are not the same thing. That is why the metric is never read alone.
The effectiveness computation: which comments exist, which were followed by a code change in the PR, and the aggregations by reviewer, team, and period. Running it twice over the same history produces the same number.
Only the classification of the comment TYPE (nit, logic, security, style). The model never decides whether a comment was effective, and reclassifying types does not change the effectiveness number.
The baseline is not defined by title, seniority, or nomination. It is the reviewers whose comments convert into changes most often, over a minimum volume of activity, recalculated as history grows. The criterion is independent of rank: nobody joins the baseline by authority, and the ruler is not tunable per contributor, which includes Releezy’s own agents.
Limits we know about and account for in the reading:
The first-party numbers quoted on this site come from customer production environments measured by Releezy Guardian. The base of the main case: 13,784 pull requests, a team of 102 developers, 11 months of history. Every number travels with the base that produced it, and whoever is measured sees their own data.
Every metric carries the history that produced it: the comments, the PRs, the time window. A recommendation from Releezy Concierge arrives with that evidence attached, and disagreement is settled by looking at the record. If you find a case where the rule counts wrong, we want to see it: the methodology improves by being contested, not by being defended.
Give engineering an ally that turns evidence into action. Improve continuously with trust, connecting every recommendation to facts and every action to its result.