We publish our numbers. Every figure on this page is measured against real, community-labeled decks, not hand-picked showcases. Each block is dated, and the rating engine keeps moving between measurements, so read the date as part of the number.
of ratings land within one bracket of the community's label
Measured on 3,585 labeled Commander decks, as of 2026-08-03
On tournament cEDH lists, where the label is beyond dispute, exact-match is 95.3% and every rating lands within one.
When our rating disagrees with a deck's label, it almost always reads the deck as stronger: 33.4% of decks rate above their label versus 10% below. Before calling that an error rate, consider the direction. Commander players are famous for underselling their own decks - "it's just a 2" is practically a table ritual - and the official bracket definitions do not bend to modesty. A deck holding a two-card infinite combo, four or more Game Changers, mass land denial, or chained extra turns has a minimum bracket under WotC's own chart, whatever its owner calls it. We enforce those definitions; self-ratings often skip them. We measured it: of the 1,196 decks we rate above their label, 21.5% carry contents whose official minimum bracket already exceeds that label - those labels are not possible under the chart. We publish the miss anyway, because the remainder is real and we would rather measure it than hide it.
Not all labels are equally trustworthy, so here is every cohort, including the ones that make us look worse. Accuracy tracks label quality almost perfectly - which is what you should expect from a rater that is measuring something real.
Decks from competitive results. The label is beyond dispute.
Players who saw our rating and confirmed it. Anchored, so read gently.
All labeled bracket 2, dating from before WotC decoupled precons from any single bracket. 19 of these 142 ship contents whose official floor is 3 or higher - even the box sandbags sometimes.
Owners rating their own decks. This is where sandbagging lives.
Smaller label sources that fit no bucket above.
Every deck here is one someone disagreed with, by definition. Low exact-match is the point of the cohort.
The ground truth here is human judgment: bracket labels from player feedback, partner-site corrections, tournament lists, and official preconstructed decks. Human labels are noisy in a particular direction - self-ratings skew low, and disagreement clicks over-represent people who felt over-rated. The same deck gets called a 2 by one table and a 3 by the next. Against that baseline, an exact-match score can never approach 100% for any rater, human or machine, and a rater that matched sandbagged labels perfectly would be reproducing the sandbagging. That is why the honest headline is the within-one number, with exact-match shown plainly above, and why the cleanest read of raw accuracy is the cohort where labels are beyond dispute: tournament cEDH lists. Decks with a partner commander are rated on the primary commander in this corpus, so partner-pairing effects are not fully captured.
| Bracket | Decks | Exact | Within one |
|---|---|---|---|
| 1Exhibition | 49 | 14.3% | 57.1% |
| 2Core | 852 | 55.9% | 93.7% |
| 3Upgraded | 1,810 | 51.8% | 99.1% |
| 4Optimized | 633 | 64.5% | 96.1% |
| 5cEDH | 241 | 83.8% | 98.3% |
Bracket 1 (Exhibition) is rare by design - decks must show active evidence of being unfocused - so its sample is small and its labels are the noisiest in the corpus.
CommanderBracket does not count cards against a checklist. The engine models how fast a deck actually threatens to win, then maps that speed to the official WotC bracket definitions:
Our rules assistant is graded the same way: against the comprehensive rules themselves, not against vibes. On a 150-question difficulty-stratified eval (June 2026), it answered 96.5% of questions correctly with zero invented rules - every wrong answer cited a real rule and reasoned incorrectly about a hard interaction, and the assistant abstains with a "ask your local judge" warning rather than guess when it is unsure.
That eval predates our current rules model, and re-grading found a handful of questions where our own answer key was wrong, which the scoring counted in our favour. Treat the percentage as approximate. The zero-invented-rules result held across every question and is the part we stand behind. Unofficial rules guidance, not a tournament ruling. For tournament-level answers, consult a certified judge.
Or run any list through the free Commander bracket checker - pasted decklists and deck-site links both work.