Accuracy

We publish our numbers. Every figure on this page is measured against real, community-labeled decks, not hand-picked showcases. Each block is dated, and the rating engine keeps moving between measurements, so read the date as part of the number.

96.6%

of ratings land within one bracket of the community's label

Measured on 3,585 labeled Commander decks, as of 2026-08-03

On tournament cEDH lists, where the label is beyond dispute, exact-match is 95.3% and every rating lands within one.

The honest breakdown

56.7%
Exact match
40%
Off by one
3.4%
Off by two or more

When our rating disagrees with a deck's label, it almost always reads the deck as stronger: 33.4% of decks rate above their label versus 10% below. Before calling that an error rate, consider the direction. Commander players are famous for underselling their own decks - "it's just a 2" is practically a table ritual - and the official bracket definitions do not bend to modesty. A deck holding a two-card infinite combo, four or more Game Changers, mass land denial, or chained extra turns has a minimum bracket under WotC's own chart, whatever its owner calls it. We enforce those definitions; self-ratings often skip them. We measured it: of the 1,196 decks we rate above their label, 21.5% carry contents whose official minimum bracket already exceeds that label - those labels are not possible under the chart. We publish the miss anyway, because the remainder is real and we would rather measure it than hide it.

By label source

Not all labels are equally trustworthy, so here is every cohort, including the ones that make us look worse. Accuracy tracks label quality almost perfectly - which is what you should expect from a rater that is measuring something real.

Tournament cEDH lists64 decks
95.3% exact100% within one

Decks from competitive results. The label is beyond dispute.

Ratings users agreed with971 decks
81.3% exact99.8% within one

Players who saw our rating and confirmed it. Anchored, so read gently.

Official preconstructed decks142 decks
75.4% exact98.6% within one

All labeled bracket 2, dating from before WotC decoupled precons from any single bracket. 19 of these 142 ship contents whose official floor is 3 or higher - even the box sandbags sometimes.

Self-rated community decks1,111 decks
57.7% exact94.4% within one

Owners rating their own decks. This is where sandbagging lives.

Other sources100 decks
70% exact98% within one

Smaller label sources that fit no bucket above.

Disputed ratings1,197 decks
30.3% exact95.6% within one

Every deck here is one someone disagreed with, by definition. Low exact-match is the point of the cohort.

How to read these numbers

The ground truth here is human judgment: bracket labels from player feedback, partner-site corrections, tournament lists, and official preconstructed decks. Human labels are noisy in a particular direction - self-ratings skew low, and disagreement clicks over-represent people who felt over-rated. The same deck gets called a 2 by one table and a 3 by the next. Against that baseline, an exact-match score can never approach 100% for any rater, human or machine, and a rater that matched sandbagged labels perfectly would be reproducing the sandbagging. That is why the honest headline is the within-one number, with exact-match shown plainly above, and why the cleanest read of raw accuracy is the cohort where labels are beyond dispute: tournament cEDH lists. Decks with a partner commander are rated on the primary commander in this corpus, so partner-pairing effects are not fully captured.

By bracket

BracketDecksExactWithin one
1Exhibition4914.3%57.1%
2Core85255.9%93.7%
3Upgraded1,81051.8%99.1%
4Optimized63364.5%96.1%
5cEDH24183.8%98.3%

Bracket 1 (Exhibition) is rare by design - decks must show active evidence of being unfocused - so its sample is small and its labels are the noisiest in the corpus.

How the rating works

CommanderBracket does not count cards against a checklist. The engine models how fast a deck actually threatens to win, then maps that speed to the official WotC bracket definitions:

  • Win-turn modeling. Ratings derive from an estimate of the turn a deck can credibly threaten a win, calibrated against thousands of real labeled decks. Two decks with the same staples can land in different brackets because they play at different speeds.
  • WotC bracket rules enforced. Game Changer counts, mass land denial, chained extra turns, and two-card combos apply the official bracket restrictions as hard floors, exactly as WotC defines them.
  • Prerequisite-aware combo detection. Combos are checked against a 40,000+ combo database including their setup requirements. A pairing that needs a five-creature board before it does anything is not treated like a turn-one win.
  • Reasons on every rating. Every bracket ships with the factors that produced it, so you can see why - and push back when you disagree. Those disagreements feed the next calibration pass.

Ask the Judge

Our rules assistant is graded the same way: against the comprehensive rules themselves, not against vibes. On a 150-question difficulty-stratified eval (June 2026), it answered 96.5% of questions correctly with zero invented rules - every wrong answer cited a real rule and reasoned incorrectly about a hard interaction, and the assistant abstains with a "ask your local judge" warning rather than guess when it is unsure.

100%
Basic
~95%
Intermediate
~94%
Advanced

That eval predates our current rules model, and re-grading found a handful of questions where our own answer key was wrong, which the scoring counted in our favour. Treat the percentage as approximate. The zero-invented-rules result held across every question and is the part we stand behind. Unofficial rules guidance, not a tournament ruling. For tournament-level answers, consult a certified judge.

Rate your deck

Or run any list through the free Commander bracket checker - pasted decklists and deck-site links both work.