The Twenty-Game Audit: Finding Your Pattern, Not Your Worst Move
There is a ritual most improving players perform after a loss. Paste the PGN into Lichess, hit “Request a computer analysis,” wait eleven seconds, and stare at the red exclamation marks. The engine says you dropped 340 centipawns on move 23. You nod grimly, resolve to “calculate better,” and close the tab.
That ritual is useless. Not slightly useless: actively counterproductive, because it feels like work and produces nothing you can train.
The problem is the sample size. A single game is one data point drawn from an enormous distribution of your possible performances. You might have been tired. Your opponent might have played an opening you’d never seen. You might have had eight seconds on the clock when you hung the rook. Stockfish 17, running at depth 22 on Lichess’s cloud, will find something wrong in every game you have ever played, including the ones you won 1-0 in 24 moves. An engine’s job is to find the best move. Your job is to find your pattern. These are different jobs and they need different methods.
Why one game lies to you
Consider two players, both rated 1450 on Chess.com rapid. Both lose a game where the engine flags a blunder on move 31, hanging a knight in a winning endgame.
Player A has hung a piece in 3 of the last 20 games, all of them after move 30, all of them with under two minutes on the clock. Player B has hung one piece in 40 games and it happened at 11pm after a double shift.
Same game, same engine output, completely different diagnosis. Player A has a time-management failure that consistently converts winning endgames into losses. Player B has nothing to fix. If both read only the single-game report, both draw the same conclusion, and one of them wastes three months doing tactics puzzles that were never the problem.
The engine cannot tell you which player you are. Only the aggregate can.
Building the dataset
Here is the actual procedure. It takes about 90 minutes the first time.
Step one: pick the games, not the losses. Take your last 20 classical or rapid games at your normal time control. Include wins. This matters more than it sounds: players hide their worst structural habits inside wins, because a win feels like confirmation. If you only audit losses you will systematically under-count the mistakes that your opponents happened to punish less.
Exclude bullet. Bullet errors are motor errors and they belong to a different study.
Step two: get them out. On Lichess, go to your profile, click the gear icon, and “Export games.” You can filter by time control and date. It hands you a single PGN file. On Chess.com, Settings → Download Games gives you a monthly archive. If you play on both, you now have two files, which is fine.
Step three: analyse them all, consistently. Lichess’s “Request a computer analysis” is free and runs Stockfish server-side at a fixed depth. That fixed depth is the point: it makes games comparable to each other. If you analyse game 3 at depth 18 in your browser and game 14 at depth 30 overnight on a desktop, your blunder counts are not measuring the same thing.
If you want more control, install the free Lichess lichess-bot-adjacent tooling or just use the Python library python-chess with a local Stockfish binary:
import chess, chess.pgn, chess.engine
engine = chess.engine.SimpleEngine.popen_uci("stockfish")
LIMIT = chess.engine.Limit(depth=20)
with open("mygames.pgn") as f:
while (game := chess.pgn.read_game(f)):
board = game.board()
prev = None
for ply, move in enumerate(game.mainline_moves()):
info = engine.analyse(board, LIMIT)
score = info["score"].white().score(mate_score=10000)
if prev is not None:
loss = abs(score - prev)
if loss >= 100:
print(game.headers["Site"], ply//2 + 1,
board.san(move), loss, board.fen())
prev = score
board.push(move)
Twenty games at depth 20 takes roughly 25 minutes on a modern laptop. Redirect that output to a text file. That file is your audit.
The spreadsheet is the whole method
Open a spreadsheet. One row per error of 100 centipawns or more, which in practice means 60 to 140 rows across 20 games at club level. Columns:
| Field | Values |
|---|---|
| Game ID | Lichess/Chess.com URL |
| Move number | integer |
| Phase | opening (≤12), middlegame (13–30), endgame (31+) |
| CP loss | integer |
| Clock remaining | mm:ss |
| Material | balanced / up / down |
| Pawn structure | IQP, closed centre, opposite castling, symmetrical, etc. |
| Error type | hung piece, missed tactic, positional drift, bad trade, king safety, time scramble |
| Engine’s move | SAN |
The “error type” column is where the judgement lives, and you will get it wrong on individual rows. That’s acceptable. With 100 rows, misclassifying eight of them does not move the aggregate.
One rule: fill in the error-type column before you look at the engine’s suggested move. Otherwise you’ll rationalise. Engine output is extraordinarily good at making you believe you “almost saw that.”
Reading the output
Now sort and count. Here’s a real-shaped example from an audit of a 1560-rated player’s last 20 rapid games (15+10):
Errors ≥100cp by phase:
Opening (≤12): 9 (7.8%)
Middlegame(13-30): 41 (35.7%)
Endgame (31+): 65 (56.5%)
Errors ≥300cp by clock remaining:
> 5:00 left: 4
2:00–5:00: 6
< 2:00 left: 19
Errors ≥100cp by structure (middlegame only):
Opposite-side castling: 22 of 41 (in only 6 of 20 games)
Closed centre: 5
Symmetrical/other: 14
Three numbers jump out and they point at the same thing from different angles. Over half the errors are in the endgame. Nineteen of the 29 serious blunders happen under two minutes. And opposite-side castling games, which are 30% of the sample, generate 54% of the middlegame errors.
That is a diagnosis. It is not “calculate better.” It is: this player reaches good positions, spends too long in the middlegame, and loses them in time trouble, and the time trouble is concentrated in sharp races they don’t know how to play by feel.
The training plan writes itself. Study three opposite-castling attacking structures until the plans are automatic, which buys back clock. Do 30 minutes of king-and-pawn plus rook-endgame drills so the last phase costs less thought. Set a hard rule: at move 20, if more than half the clock is gone, play the obvious move.
Compare that to what the single-game report said on any one of those 20 games. It said: “Blunder. 24…Rxd4 was better.” True and worthless.
The patterns you should expect to find
Certain clusters recur at club level, and knowing their shape speeds up recognition.
The 3-to-5 opening hole. All 20 games pass through the same four openings. If 9 of your opening errors happen in one variation, that’s not a calculation problem, it’s twenty minutes of database work. Open the Lichess opening explorer, filter to your rating band, and look at what the 1400–1600 pool actually plays against your setup. Your engine report will label the move an inaccuracy worth 60cp; the explorer will tell you 38% of opponents in your band play it, which is the number that determines whether it’s worth learning.
The conversion collapse. Count games where you were +2.0 or better at some point and did not win. If that number is 4 or more out of 20, technique is your single highest-leverage study area and nothing else comes close. This one is invisible in single-game analysis because the engine happily reports “you were winning” without noting that you are always winning and never converting.
Evaluation illiteracy. Look for positions where the engine says +0.3 and you played the aggressive move that gives it back. A +0.3 is not an attack. It is a very slightly better pawn structure. Club players routinely torch small advantages trying to make them big, and the audit shows it as a specific move pattern: sacrifices and pawn pushes launched from evaluations under +0.6. If you see six of those, you have learned to read an engine number, which is worth more than any opening you could memorise. The pillar piece on analysing your own games with an engine covers what the evaluation bar is actually telling you in more depth; read it alongside your own spreadsheet rather than on its own.
Colour asymmetry. Split all your counts by White and Black. A 100-point performance gap between colours is common and almost always traces to one specific defensive setup you don’t trust.
What not to do with the data
Don’t chase the single worst move in the dataset. The -740cp catastrophe on move 34 of game 11 feels like the most important thing in the file. It is the least important: it happened once, under conditions that produced it once.
Don’t act on a count below 4. Three occurrences of anything in 20 games is noise. With roughly 100 error rows, a pattern needs to clear 8-10% of your total to be worth restructuring your study around.
And don’t re-audit next week. Twenty games at 15+10 is maybe six weeks of normal club play. Fix the one thing your audit identified, play through that many games again, and then re-run the same procedure with the same depth setting. The second audit’s value is the difference: opposite-castling errors went from 22 to 7, endgame share dropped from 56% to 38%, but blunders under two minutes barely moved. Now you know your opening work landed and your clock discipline didn’t.
That comparison is the only honest measure of whether the last six weeks of study did anything at all, and it’s the reason to keep the spreadsheet rather than the engine report. Your rating will move for a hundred reasons, half of them your opponents’. The error distribution moves only when you change.