Can An LLM Coach Your Chess? Establishing A Baseline
Ask a language model to find the best move in a sharp position and you will get a confident answer, a plausible-sounding justification, and, often enough, a move that is not legal. Ask the same model to explain why Stockfish prefers 14.Rfe1 over 14.Rae1, with the engine’s lines pasted in front of it, and you get something a coach charging £60 an hour would recognise as useful. Same model. Same session. The difference is what job you handed it.
So: can AI coach you at chess? Not on its own, and not in the way the marketing implies. The useful answer is narrower and more actionable. Language models cannot calculate. They can explain, plan, prioritise and summarise. An engine can calculate and cannot do any of the other four. Bolt them together in that order and you have something worth your Tuesday evenings. Use either alone and you have a toy.
The thing the model physically cannot do
Stockfish 17 sits north of 3600 on the CCRL 40/15 list. It gets there by searching: tens of millions of nodes per second, alpha-beta pruning, an NNUE evaluation network scoring leaf positions. There is a tree, and it walks the tree.
A language model does none of that. It produces the next token given the previous ones. When it writes “15.Nxe6 fxe6 16.Qxe6+ Kh8 17.Qxe7 wins the exchange,” it is not verifying that the knight can reach e6, that the pawn can capture, or that the queen has a path. It is producing text that pattern-matches to annotated games it has seen. Sometimes the pattern is right. In a quiet Carlsbad structure it is right surprisingly often, because the moves are thematic and the corpus is full of them. In a tactical melee with three pieces hanging it falls apart, and it falls apart while sounding exactly as confident as when it is correct. That last part is the whole problem.
The baseline: thirty positions, three buckets
Before you trust any of this, measure it. The protocol below takes about ninety minutes and you only need to run it once. Numbers quoted here come from one run in September 2026 against the current chat models; yours will differ, and that is fine, because the shape of the result is what matters.
Pull thirty positions from your own Lichess or Chess.com games, ten per bucket:
- Tactical: a forcing sequence exists, three to six plies deep.
- Positional: no tactics, engine eval between -0.4 and +0.4, a structural decision on the board.
- Endgame: five to eight pieces, technique required.
For each position, run two trials. Trial A gives the model the FEN and nothing else: “Find the best move and explain.” Trial B gives the model the FEN plus Stockfish output at depth 22, including the top three moves with evaluations and principal variations, and asks it to explain the engine’s preference in plain language, then name the one idea to remember.
Score four things. Legality of any move the model names. Agreement with the engine’s top move. Whether the explanation contains a factual claim about the position that is false (a piece on the wrong square, a pawn that cannot advance, a mate that isn’t there). And whether the takeaway is specific enough to train.
In our run, Trial A produced an illegal or non-existent move in 11 of 30 positions, with nine of those eleven in the tactical bucket. Trial A’s endgame answers were worse than its positional ones, which surprised nobody who has watched a model try to count tempi in a pawn race. Trial B produced zero illegal moves, because it was never asked to generate one, and the false-claim rate dropped from roughly a third of answers to two out of thirty. Both remaining errors were of the same type: the model over-generalised the engine’s line into a rule that didn’t hold.
That gap is your baseline. Everything else in this post is downstream of it.
A worked example, both ways
Take the Blackburne Shilling Gambit, a line club players actually meet. After 1.e4 e5 2.Nf3 Nc6 3.Bc4 Nd4 White grabs with 4.Nxe5, and Black plays 4…Qg5.
Handed the raw position after 4.Nxe5 and asked what Black should do, one model told us Black should play 4…Nxc2+ “forking king and rook,” which is not a check and does not fork anything useful. Another found 4…Qg5 but described the follow-up as winning a pawn, missing that 5.Nxf7 Qxg2 6.Rf1 Qxe4+ 7.Be2 Nf3 is mate.
Now give it Stockfish’s output instead:
info depth 24 seldepth 32 score mate 4 pv g8g5 e5f7 g5g2 h1f1 g2e4 f1e1 e4e2
bestmove d8g5
With that in front of it, the model wrote back that White’s 4.Nxe5 loses because the knight on e5 abandons g2 while the knight on d4 covers f3 and e2, that Black’s queen infiltrates on g2 and the mate is delivered by the knight rather than the queen, and that the pattern to remember is “when your opponent’s knight leaves the kingside, count the defenders of g2 before you count material.” That is a coaching sentence. The engine will never write it.
Centipawns, and why the numbers mislead you
Most improvers read +1.2 as “winning” and 0.00 as “boring.” Both readings cost rating points.
| Engine eval | What it actually means at 1200-1800 |
|---|---|
| 0.00 to ±0.30 | Statistically level. Whoever understands the structure better will win. |
| ±0.30 to ±0.70 | A real but small edge. Below 1800 this converts maybe 55% of the time. |
| ±0.70 to ±1.50 | Clear advantage. Needs a plan, not a tactic. |
| ±1.50 to ±3.00 | Winning with technique you may not yet have. |
| ±3.00 and up | Decisive unless you hang something. |
Lichess does not classify mistakes by centipawns at all, which is worth knowing. It converts the evaluation into a win percentage and flags an inaccuracy at a drop of roughly 10 percentage points, a mistake at 20, a blunder at 30. That is why a move that swings the eval from +6.0 to +3.5 gets no flag while one that goes from +0.3 to -0.4 gets called a mistake. Feed a model your accuracy report without that context and it will happily invent a narrative about your “collapse in winning positions” from noise.
Average centipawn loss is the other number people misread. A strong club player runs somewhere around 35 to 50 ACPL in classical games; around 1200 you are more often looking at 70 to 100. One game’s ACPL tells you almost nothing because a single blunder dominates the mean. Twenty games’ worth, segmented by phase, tells you a great deal, and segmenting twenty games by phase is exactly the sort of tedious summarisation a language model is genuinely good at.
The configuration that survives contact
Here is the pipeline that held up across every test we threw at it. Engine as sensor. Model as narrator and planner. You as the person who decides what to work on.
Export twenty of your recent games as PGN from Lichess (the export includes evaluations if you ran the computer analysis) or Chess.com. Run them through a GUI that gives you per-move eval: En Croissant and Nibbler both do this well, and Nibbler will happily drive Leela alongside Stockfish if you want a second opinion on quiet positions. Then hand the annotated PGN to the model with a prompt that forbids it from doing the one thing it cannot do.
You are analysing my games. You have Stockfish evaluations for every move.
RULES:
- Never propose a move the engine did not list. If you want to discuss a
move, quote its eval from the PGN.
- Never claim a tactic exists unless it appears in a principal variation.
- If you cannot support a claim from the data, say "not visible in this data".
TASKS:
1. Group my blunders (win% drop >= 30) by cause: hanging piece, missed
opponent threat, wrong plan, time pressure, endgame technique.
2. Report the count and the average move number for each group.
3. Name the single most frequent cause and quote three examples.
4. Give me one week of training: four sessions, 45 minutes each, each
targeting that cause specifically. Name the exact Lichess or Chessable
resource for each session.
The first three tasks are counting and grouping. The fourth is planning. None of them require the model to see a chessboard. When we ran this on a 1450-rated player’s twenty games, the output was: 23 blunders, 14 of them classified as “missed opponent threat,” average move number 21.4, clustering hard in the transition out of the opening. The plan that came back was four sessions of two-move tactics with a deliberate blunder-check habit attached, plus a rule about writing down the opponent’s threat before every move in the 15 to 25 range. That is a coach’s diagnosis, produced from data the coach didn’t have to gather.
If you want the specific model and tool comparisons behind this, including which chat interfaces will accept a 20-game PGN without truncating it and which prompt structures produced the fewest hallucinated claims, that testing lives in AI Coaching Tools And LLM Chess Prompts, Tested.
What this does not fix
Your model still has no idea how strong you are. Ask it to play at 1400 and it will play a mixture of grandmaster moves and beginner blunders, because “1400-strength chess” is not a thing it can simulate. For that, Maia exists: neural networks trained on Lichess games at specific rating bands from 1100 to 1900, matching the human move just over half the time, which is considerably better than Stockfish at guessing what a human will actually play. Maia predicts. It does not explain. You are back to pairing.
Nor does any of this survive you skipping the engine step. The moment you paste a bare FEN and ask “what should I play here,” you are back in Trial A, and Trial A got a move wrong roughly a third of the time while sounding completely certain. The discipline is the product.
Run the thirty positions this week. Keep the scorecard somewhere you will see it in three months, because the models will change and your baseline is the only thing that will tell you whether they changed in a direction that helps you.