AI Chess

Inaccuracy, Mistake, Blunder: What The Labels Actually Measure

Open any game on Lichess, click “Request a computer analysis”, and thirty seconds later you have a verdict: 3 inaccuracies, 2 mistakes, 1 blunder. The words feel like judgement. They sound like a coach who watched you play and is now telling you, gently but firmly, where you went wrong.

They aren’t. They’re arithmetic.

Every one of those labels comes from a single number: how much your evaluation dropped between the position before your move and the position after it. No consideration of what kind of error it was, whether you had a plan, whether the position was sharp or dead quiet, whether a grandmaster would have found the right move or whether it required six moves of calculation. Just a subtraction, compared against a threshold somebody picked.

The actual cut-offs

Here is what Lichess uses. The numbers are in “win percentage” rather than raw centipawns, which matters, and I’ll come back to why.

LabelWin % drop
Inaccuracy5 or more
Mistake10 or more
Blunder20 or more

Lichess converts the engine’s centipawn score into an expected win percentage first, using a sigmoid: 50 + 50 * (2 / (1 + exp(-0.00368208 * centipawns)) - 1). Then it compares the win percentage before your move with the win percentage after.

Chess.com does it differently. Its Game Review runs on raw centipawn loss with thresholds that shift depending on how good the position already was, plus the “Best / Excellent / Good” tiers on the positive side and the Brilliant and Great Move labels layered on top. The exact formulas aren’t published and have changed at least twice since 2021. What’s consistent across both sites is the shape of the thing: a scalar drop, bucketed.

So consider two positions.

Position A. You’re equal, 0.0. You play a move and the engine now says -0.45 from your side. Win percentage went from 50% to roughly 46%. Drop of 4 points. Lichess says nothing at all. Clean sheet.

Position B. You’re winning by +6.0 (win percentage ~98.9%). You play a sloppy move and it’s now +3.0 (win percentage ~96.6%). You threw away three whole pawns of advantage and Lichess also says nothing, because the drop is 2.3 points. Two positions, wildly different in chess terms, same silence from the analysis.

That sigmoid is doing a lot of work, and it’s mostly correct work. Going from +6 to +3 genuinely doesn’t change the result of the game if you have hands. Going from 0.0 to -0.45 in a sharp middlegame might well cost you the point. But it also means the label density in your game reports is heavily concentrated in the range where the eval sits between roughly -2 and +2, and almost absent outside it. A player who wins by converting a +8 into a +4 into a +2 into mate looks flawless in the report. A player who ground out a level endgame and lost it looks like a disaster.

Where this genuinely misleads you

The problem isn’t that the thresholds are wrong. It’s that a single scalar drop cannot distinguish between two completely different failures, and your training plan depends entirely on which one you committed.

Failure type one: the one-move oversight. You had a good position. You didn’t see that the knight was hanging, or that Qh4 was mate, or that after your capture there’s a check that wins the piece back with interest. Eval goes from +0.6 to -4.2. Win percentage drops ~45 points. Lichess: Blunder.

Failure type two: the slow drift. You played eleven moves in a row that were each individually fine-ish, each losing between 3 and 6 win percentage points, and at the end of it you have a bad bishop, no squares for your knight, a weak c-pawn and a structurally lost position. Eval went from +0.6 to -2.8 over eleven moves. Lichess: two inaccuracies, maybe one mistake, seven moves with no label at all.

The oversight gets flagged in red and screams at you. The drift produces a report that looks mildly untidy. And the drift is by far the more important thing to fix, because it reflects something you don’t understand, whereas the oversight reflects something you didn’t look at.

Here’s the asymmetry that actually costs rating points. Blindness to a hanging piece is a scanning habit. You fix it with a physical routine: before you commit, check every piece of yours that your opponent’s pieces can reach. It’s cheap, it’s trainable, and it stops happening once you build the habit. Positional drift is a knowledge gap. It takes months, and it takes studying structures, not puzzles. The report that flags the cheap fix in red and buries the expensive one in grey is actively pointing you at the wrong work.

Watch it happen

Take this line from a real 1400-level game (colours anonymised, White to move after 1.e4 e5 2.Nf3 Nc6 3.Bb5 a6 4.Ba4 b5 5.Bb3 Nf6 6.d3 Bc5). Here’s what running the position through Stockfish 16 at depth 22 in the Lichess board editor gives you, move by move, if White continues with a sequence of slightly passive but never catastrophic choices:

Move       Eval      Win%     Drop    Label
7.Nc3      +0.31     54.5     0.0     -
8.O-O      +0.28     54.1     —       (opponent's move between)
9.h3       +0.09     51.3     2.8     -
10.Nd5     -0.14     48.0     3.3     -
11.Nxf6+   -0.38     44.5     3.5     -
12.Be3     -0.61     41.2     3.3     -
13.Bxc5    -0.85     37.8     3.4     -
14.Qd2     -1.20     33.0     4.8     -
15.Rad1    -1.56     28.3     4.7     -

Eight moves. Not a single label. White has gone from a comfortable +0.31 to -1.56, which is the difference between “fine” and “probably losing this against anyone competent”. Every drop landed between 2.8 and 4.8, all of them under the 5-point inaccuracy floor.

Now the same game, move 24, where White drops a rook to a two-move tactic: eval goes -1.56 to -6.90, win percentage from 28.3 to 8.0, drop of 20.3. Blunder. Red flag. That’s the move the report tells you to study.

But the rook was already in trouble by move 15. The blunder on move 24 was made possible by the eight moves nobody flagged. If you take the report at face value you’ll spend your next session on tactics puzzles when what you needed was twenty minutes with a book on the Spanish and the question “why does h3 followed by Nd5 hand over the initiative?”.

What to do instead

Read the eval graph, not the label list. On Lichess the graph sits under the board and it’s the single most useful thing in the analysis. You’re looking for shape. A cliff means an oversight: one move, big drop, trainable with habit work. A staircase means drift: a long diagonal slide with no single dramatic step, and that’s a knowledge problem. If your graphs are mostly staircases you are not a tactically weak player who needs more puzzles, whatever your accuracy score says.

Compute your own threshold. Pull the centipawn-loss numbers yourself and look at the distribution rather than the buckets. Lichess exports a PGN with [%eval ...] comments on every move when you download an analysed game. A short Python script over a few hundred of your games, counting how many moves fall in each 2-point win-percentage band, tells you more than any accuracy figure. If you have a fat lump of moves in the 3-to-5 band, you have a drift problem that the site has never once mentioned to you.

Check whether the drop is even real. Lichess’s server analysis runs a shallow, fast search: Stockfish at depth 20-ish with very little time per move. Chess.com’s default Game Review is comparable. Both will call something an inaccuracy that a deeper search rates as fine, and both will occasionally miss a genuine error because the refutation was at depth 28. Before you build a training plan around a flagged move, re-run that position in the local analysis board and watch the eval as depth climbs. If the number wobbles between +0.2 and -0.4 as depth goes from 18 to 30, the engine is telling you the position is unclear, not that you erred. This is the whole reason it’s worth understanding how engine evaluations are produced and where they lie to you before you treat any of these numbers as fact.

Ask a different question about each flagged move. For every move labelled mistake or blunder, ask: could I have found this with the method I already use? If yes, it’s an execution failure, and execution failures respond to routine and time management. If no, if the refutation was five moves deep and required seeing a quiet intermediate move, it’s a calculation-depth problem and it responds to a completely different kind of practice. The label doesn’t distinguish these. You have to.

The thresholds aren’t sacred

Lichess’s 5/10/20 split is a reasonable choice by developers who needed some number to draw a line at. It’s in the open-source code, you can go read it, and there’s no theory of chess error behind it. Chess.com’s version is tuned differently and produces different counts on the identical game: I’ve seen the same 40-move game come back as 1 blunder, 2 mistakes, 4 inaccuracies on Lichess and 2 blunders, 1 mistake, 7 inaccuracies on Chess.com. Neither is lying. They just drew their lines in different places.

Once you know the cut-offs are arbitrary, the labels stop being a report card and become what they always were: a rough index into a list of positions worth looking at. Some of the entries are wrong. Some of the most important moves in the game aren’t in the index at all. The person who has to decide which is which is you, sitting with the board, asking why a move felt right at the time.

Start with your last ten losses. Look at the graph shape before you look at the label count, and see which of your games were cliffs and which were staircases. The split will probably surprise you.