How rankings work

Battle Rating measures strength—and shows what we still do not know.

Every deck receives a familiar integer rating, a best-estimate rank, and likely ranges. The simulator then spends its next game on the matchup expected to make the ranking more accurate.

What the number means

A deck’s field score is its estimated average match score against one uniformly selected other deck in the same pool, with a draw worth half a win and both play/draw orientations averaged. We convert that score to Battle Rating on a log-odds scale: 1600 means a 50% field score, about 1791 means 75%, and about 1409 means 25%.

The number is relative to its rating universe. A Daily Draft universe contains that day’s canonical candle decks and every submitted human deck. Ratings from different days or different pools should not be compared directly because their fields are different.

Read the range, not only the point

“Battle Rating 1684, likely 1590–1772” means 1684 is our current best display value and the second pair is a 90% posterior credible interval under the model. A rank range works the same way. Wide or overlapping ranges are an honest sign that more simulation is needed; they are not hidden behind an overconfident leaderboard.

Early ratings are pulled toward 1600. As evidence arrives, the estimate can move and its range usually narrows. A rating change after someone else’s game is normal: every result helps explain the strength of previous opponents and the field as a whole.

Playing more games does not itself earn rating points. Match count and the width of the likely range are not part of the displayed rating formula. More evidence can move the strength estimate up or down, but merely narrowing an unchanged estimate leaves its Battle Rating unchanged.

Matchups and play/draw are modeled

The model learns a global strength for each deck, a shared advantage for the first-listed side, and a small, strongly regularized effect for each observed matchup. This lets it represent rock-paper-scissors fields without pretending every result is explained by a single skill number. The public rating remains coherent because it averages each deck’s learned matchups over the complete field.

How we spend simulation games

For each eligible next matchup, the scheduler considers a win, draw, and loss, updates the hypothetical posterior, and estimates how many deck pairs the displayed order would get wrong. It chooses the pairing block with the greatest expected reduction in that error. Pending games are kept on different decks when possible, and play/draw orientation is balanced.

In precise terms, this is a one-step Bayesian experimental-design policy: among the shortlisted eligible matchups, under the current model and equal game costs, it minimizes expected posterior Kendall inversions after the next result. It is locally optimal in that stated sense—not a claim that an approximate model can find a globally perfect schedule for every possible metagame.

Daily Draft’s promised standard-candle games are a product constraint, so they are played even when the scheduler might choose another pairing. They still become evidence in the same model. Additional Arena games are scheduled in compact pairing blocks so Forge gets good throughput, while every completed block can still influence the next scheduling choice.

Important limits

  • Forge measures how these decks are played by its configured AI, not perfect human play.
  • A credible interval is conditional on the model; bugs or systematic simulator bias are outside it.
  • A field-score rating summarizes a matchup matrix. Open the matchup record when the specific opponent matters.
  • More games improve estimates, but no finite simulation can make a close ordering certain.