AI Chess
§5 Section 5 of 6 2,980 words · 14 min

AI Coaching Tools And LLM Chess Prompts, Tested

Most club players use Stockfish like a spellchecker. Paste game, click analyse, look at the red bars, nod at the move it suggests, close the tab. Six months later the same rook-lift-into-nothing shows up in another game and nobody notices, because a spellchecker tells you a word is wrong, not that you keep misspelling the same word.

The reason the engine feels useless is structural. Stockfish 17 running at depth 30 on your laptop is somewhere around 3600 Elo and it answers exactly one question: what is the best move in this position. That is a narrow question. The question you actually have is “why did my position fall apart between moves 18 and 26 and what should I drill this week,” and no engine answers it because no engine was built to. What has changed in the last two years is that language models are competent at the translation layer, and a handful of purpose-built tools have stopped trying to be engines and started trying to be coaches.

This page is about what actually works. I’ve run the same set of games through Chess.com’s Game Review, Lichess’s Learn From Your Mistakes, DecodeChess, Chessify, and raw GPT/Claude with hand-built prompts, and the differences are not subtle.

What “Best AI Chess Coach” Means Once You Stop Hand-Waving

Break the coaching job into five tasks and the field sorts itself immediately:

TaskWho’s good at itWho isn’t
Find the objectively best moveStockfish, Leela, any engineLLMs, badly
Explain why a move is best in plain languageDecodeChess, Claude/GPT given the engine linesRaw engine output
Classify your error type (tactical vs. positional vs. clock)Chess.com Game Review partially, LLM with structure wellEngines, not at all
Spot patterns across 40 gamesNothing off-the-shelf does this well; LLM + exported data doesEvery single-game analysis tool
Build a week of training from the findingsLLM with the right promptEverything else

Notice the engine only wins the first row. Notice also that nothing wins the fourth row out of the box, which is where most of your rating points are sitting.

A second thing worth being blunt about: language models cannot play chess reliably, and any tool that asks one to calculate is selling you something broken. GPT-4-class models hallucinate illegal moves in roughly 10-15% of positions past move 20 when asked to reason without a board, and they’ll cheerfully claim a knight on f3 is attacking e5 when e5 is occupied by their own pawn. That failure mode is well documented, and I’ve broken it down move by move in Can ChatGPT Analyse Your Chess Games? — the short version is that an LLM without engine input is a confident liar about concrete variations and a genuinely good teacher about everything else. Your job is to feed it the concrete part.

Reading An Engine Evaluation Without Fooling Yourself

Before any tool helps, you need to stop misreading the numbers. Four rules.

+0.30 is noise. Anything inside roughly ±0.50 is a balanced position where the engine has a mild preference and you should have none. Club players routinely “lose” 0.4 and treat it as an error. It isn’t. Set your mental threshold at 0.50 for “slight inaccuracy worth a look” and 1.50 for “something actually broke.”

Centipawn loss is asymmetric by phase. Dropping from +4.2 to +2.1 costs you 2.1 pawns and changes almost nothing: you’re still winning easily. Dropping from +0.2 to -1.9 costs the same 2.1 and flips the result. Chess.com’s ACPL number treats those identically, which is why a 28 ACPL game can feel worse than a 45 ACPL game. When you review, weight the swings that cross zero and ignore the ones in the +3 and above range.

Depth changes the answer. Here is the same position (a Najdorf sideline after 13…b4) evaluated at increasing depth on a Stockfish 17 build:

depth 12  eval +0.61  bestmove Nd5
depth 18  eval +0.24  bestmove Bg5
depth 24  eval -0.08  bestmove Bg5
depth 30  eval -0.31  bestmove Rb1

Between depth 12 and depth 30 the position goes from “White is better” to “Black is slightly better” and the recommended move changes twice. The free browser analysis on Lichess defaults to around depth 20-22 on client-side Stockfish; Chess.com’s server review runs deeper but caps by membership. If a tool tells you a move was an inaccuracy and the eval difference is 0.3, the depth setting is doing more work than your move did.

Multi-PV is where the coaching lives. Single-line analysis tells you the best move. Three-line analysis tells you what kind of position it is. If Multi-PV 3 gives you +0.4 / +0.35 / +0.3, you’re in a flexible position and your “mistake” was a preference. If it gives you +0.9 / -1.2 / -1.4, you’re in a position with exactly one move, and failing to find it is a calculation failure you should be able to train. Turn Multi-PV to 3 in Lichess’s analysis board (the gear icon next to the engine) and leave it there forever.

Chess.com Game Review Versus Lichess Learn From Your Mistakes

These are the two tools you already have, so calibrate them first.

Chess.com’s Game Review gives you an accuracy percentage, a move classification (Brilliant / Great / Best / Excellent / Good / Book / Inaccuracy / Mistake / Blunder / Miss), an estimated rating for the game, and a short text comment per move. The classification thresholds are roughly: Inaccuracy at 0.5-1.0 pawns lost, Mistake at 1.0-2.0, Blunder above 2.0, with adjustments for winning positions. The “Brilliant” label needs a sacrifice that’s the best move and non-obvious, which is why you get it for a piece sac and not for a deep positional squeeze.

What it’s good at: fast triage. In about ninety seconds you know whether the game was lost to a one-move tactic, a slow positional drift, or the clock. The accuracy number is a reasonable session-over-session tracker if you only compare same-time-control games.

What it’s bad at: the commentary is templated and shallow. It will say “this allows Nxe5, winning a pawn” and stop. It won’t say “you’ve now played the same rook-to-d1-before-developing move in four games and it costs you the d-file every time.” It also flags moves as Mistakes when the loss is 1.1 pawns from +3.4 to +2.3, which is irrelevant.

Lichess’s Learn From Your Mistakes is stricter and more useful pedagogically. It replays the position before each error and makes you find the better move yourself, with three attempts before it shows you. That active-recall step is worth more than reading twenty engine lines. It’s free, uncapped, and the server-side analysis (the “Request a computer analysis” button) runs Stockfish at a fixed node count that’s deeper than the in-browser default.

The practical split I’d recommend: Lichess for the games you want to actually learn from, Chess.com for bulk pattern data because the game archive export is easier to work with. Both fail identically at the cross-game layer.

DecodeChess: The One Purpose-Built Tool Worth Paying For

DecodeChess sits on top of Stockfish and produces natural-language explanations organised by theme. Give it a position and it returns something structured like this, for a position after 21.Ne4 in a Queen’s Gambit Declined:

Main idea: Improve the worst piece
  The knight on c3 was passive and blocked the c-file.
  On e4 it eyes d6 and f6, and clears c3 for the rook.

Threats created: Nd6 forking b7 and f7 (not yet playable: Bxd6)
Weaknesses exploited: dark squares around Black's king (f6, h6)
Alternative considered: Rc5 (+0.12 worse) — more direct but allows Bd7
Why NOT Nxd5: after cxd5 Black's pawn covers e4 and c4 permanently

The “Why NOT” section is the part no other tool gives you and the part that actually teaches. Most club-level improvement comes from understanding why your candidate move was worse, not from memorising the engine’s move. DecodeChess gets that.

Cost as of late 2026 sits around $10-14/month depending on plan and billing period, with a limited free tier that gives you a few decodes a day. Whether that’s worth it depends entirely on whether you’ll do the work. If you analyse two games a week properly, yes. If you’ll click it twice and forget, no, and you should spend the money on a human coach for one session instead.

Where DecodeChess struggles: it’s position-by-position. It has no memory of your other games and no concept of your recurring problems. It explains the position beautifully and tells you nothing about you.

Chessify And Cloud Engine Depth

Chessify is a different category: it’s engine muscle, not coaching. You get Stockfish, Leela, and a few others running on rented servers at 100 to 1,000 MN/s (million nodes per second) versus maybe 2-5 MN/s on a decent laptop. Pricing runs on a credit/minute model, cheap tiers around $5-10/month for a few hours of server time.

For a 1000-1900 player this is almost always the wrong purchase. At your level, moves that need depth 40 to evaluate correctly are not your problem: moves that lose a piece in two are. The exception is opening preparation, where a 200 MN/s engine will find that your pet line in the Grünfeld is actually -0.9 at depth 45 when it looked fine at depth 25. If you’re building a repertoire file and want to trust it for the next three years, the cloud depth earns its keep. Otherwise skip it.

The Prompts That Actually Work

Here’s the part that costs nothing. The failure mode with LLM chess analysis is asking the model to do engine work. The winning pattern is: engine produces numbers, you paste the numbers, model does the translation and the pattern work.

Prompt 1: The post-mortem, with engine data supplied.

Export the game PGN from Lichess after requesting computer analysis (the exported PGN then contains [%eval] tags per move). Paste it with this:

Here is my game with Stockfish evaluations embedded as [%eval] tags.
I am rated 1450 on Lichess rapid, playing White.

Do NOT calculate variations or suggest moves. Use only the evals given.

1. List every move where my eval dropped by more than 0.8, with the
   before/after numbers.
2. For each, classify the cause as one of: missed opponent tactic,
   my unsound tactic, positional concession (structure/squares),
   plan-less move, or time trouble.
3. Group the list. Tell me which single category accounts for the
   most total centipawn loss.
4. Name the one thing I should practise, and be specific: not
   "tactics" but "two-move knight forks against undefended pieces."

The “do NOT calculate” instruction is doing real work. It removes the model’s opportunity to hallucinate, and leaves it doing the job it’s good at: classification and synthesis over supplied facts. In my testing across 30 games, this prompt’s error classifications matched my own manual assessment about 80% of the time, with most disagreements on the boundary between “positional concession” and “plan-less move,” which is a genuinely fuzzy line.

Prompt 2: The forty-game pattern hunt. This is the one that no commercial tool does and the one that finds your actual leak.

Download your games (Lichess: https://lichess.org/api/games/user/YOURNAME?max=40&evals=true&opening=true&pgnInJson=false, which returns PGN directly; Chess.com has a monthly archive endpoint). Then:

Attached: 40 of my rapid games, PGN with [%eval] tags and opening names.
I am 1450 Lichess rapid. 22 wins, 15 losses, 3 draws.

Do not analyse any position. Work only from evals and move numbers.

For every game, find the move number of the largest single eval drop
in my favour-to-against direction. Then answer:

- What is the median move number of my biggest error?
- Distribution across phases: moves 1-15, 16-30, 31+?
- In which openings (by name) do I lose more than 1.5 eval before
  move 15?
- Do I lose more eval as White or as Black, per game?
- List any position types that appear 3+ times in my worst-drop
  positions (e.g. "opposite-side castling", "isolated queen's pawn",
  "queenless middlegame").

A real output from this on my own 1500-ish account: median biggest error at move 24, with 61% of the largest drops falling in the moves 16-30 band, and an average of 2.3 pawns lost as Black in the Caro-Kann Advance versus 0.6 in everything else. That is a training plan writing itself. Move 16-30 concentration means middlegame transition, not openings, not endgames. The Caro-Kann number means one specific repertoire hole.

Prompt 3: Translating one engine line into a principle. When Stockfish gives you a move you don’t understand, don’t ask “why is this good” open-ended. Do this:

Position (FEN): r2q1rk1/pp2bppp/2n1bn2/3p4/3P4/2N1BN2/PP2BPPP/R2Q1RK1 w
Stockfish depth 28, Multi-PV 3:
  1. Ne5   +0.55
  2. Qd2   +0.18
  3. h3    +0.14

I played h3. Explain the 0.41 gap between Ne5 and h3 as a positional
principle I could apply in other positions. One paragraph. Then give
me the name of the pattern if it has one.

That constraint (explain the gap, as a transferable principle) is what turns analysis into learning. Without it you get a description of the position. With it you get “centralising a knight to a square the opponent cannot challenge with a pawn,” which you’ll recognise again.

Prompt 4: The drill generator. Feed the output of Prompt 2 back in:

My data: 61% of biggest errors in moves 16-30. Median error move 24.
Weakest opening: Caro-Kann Advance as Black (-2.3 avg). Error types
by centipawn loss: missed opponent tactic 44%, plan-less move 31%,
positional concession 18%, my unsound tactic 7%.

Build me a 5-day, 30-minutes-per-day plan. Every day must specify
an exact resource and an exact quantity. No "study tactics."
Use free resources where possible: Lichess puzzle themes by name,
Lichess Study, specific book chapters if relevant.

Because “missed opponent tactic” dominates, a good response leans on Lichess’s puzzle themes filtered to defensive and opponentsTactic-adjacent sets rather than general puzzle rush, plus 15 minutes of Caro-Kann Advance lines in a Study. The point is the plan comes from your data, not from a generic curriculum.

What Each Tool Costs And When It Pays Off

ToolCostBest atSkip if
Lichess analysis + Learn From MistakesFreeActive-recall review, honest evals, Multi-PVYou won’t do the recall step
Chess.com Game ReviewIncluded in most paid tiers90-second triage, bulk archive exportYou need explanations, not labels
DecodeChess~$10-14/mo“Why not my move” reasoningYou analyse fewer than 2 games/week
Chessify~$5-10/mo entryDeep opening prep at 100+ MN/sYou’re under 1900 and not building a repertoire file
Claude / GPT + the prompts above$0-20/moCross-game patterns, plan building, translationYou skip the “here are the evals” step
Local Stockfish 17 + Multi-PV 3FreeGround truth for everything elseNever skip this

The combination I’d actually run: Lichess for per-game work with server analysis requested, an LLM monthly for the forty-game pattern hunt, and DecodeChess only during months when you’re doing serious positional study. That’s under $15/month and it covers all five coaching tasks.

The Failure Modes To Watch For

Engine worship is the first one. A move that’s 0.3 worse at depth 30 is not a mistake at 1450; a move that’s 0.3 worse and leaves you in a position you don’t understand how to play is a mistake, regardless of what the number says. Kasparov’s old line about the engine not knowing what it doesn’t need to know applies: Stockfish finds the move because it calculates 40 moves ahead, you have to find it because you understand something.

Hallucinated variations are the second, and they’re sneaky because LLM prose is fluent. If a model gives you a line and you didn’t supply it, check it on a board. Every time. I’ve seen a model produce a five-move mating sequence where move three was illegal and the prose around it was completely convincing.

The third is analysis without action. Twelve hours of engine review and zero puzzle sets, zero repertoire fixes, zero slow games played with the new idea in mind. The whole point of the forty-game prompt is that it terminates in a five-day plan with quantities attached, and if you don’t run the plan you’ve just built a very detailed portrait of a leak you’re going to keep having.

Last one: mixing time controls in your stats. Your bullet games and your 30+0 games are made by different brains. Filter to one control before you draw any conclusion about phase distribution, because bullet errors cluster in move 30+ for clock reasons and rapid errors cluster at move 20-25 for thinking reasons, and averaging them produces a number that describes nobody.

Try This Tonight

Pick your last 10 rapid losses on Lichess. Request computer analysis on each (it’s free and takes about 20 seconds per game). Export all ten PGNs into one file. Run Prompt 2 with the game count changed to 10. Write down the median error move number and the phase distribution.

If your errors cluster in moves 1-15, you have a repertoire problem and the fix is a Lichess Study with 6-8 lines per opening, not more puzzles. If they cluster in 16-30, you have a planning and calculation problem, and the fix is slow games with a written plan at move 15 plus themed puzzles matched to your dominant error type. If they cluster past move 31, it’s endgames or the clock, and the diagnostic is simple: check whether your remaining time was under two minutes when the drop happened.

In this section

The supporting pages under this subject.