AI Chess

Can ChatGPT Analyse Your Chess Games? A Tested Answer

Short version: yes, but not in the way most people assume, and almost never in the way the YouTube thumbnails promise. ChatGPT cannot reliably tell you whether 23.Nxe6 wins a pawn. It can tell you, with genuine insight, why you keep reaching positions where 23.Nxe6 is even a question.

I ran roughly forty of my own Lichess games through GPT-5.2, Claude Opus 5 and Gemini 3 Pro over six weeks, cross-checking every claim against Stockfish 17 running at depth 30 on a local machine. What follows is what actually happened, including the failures, with the prompts and numbers that produced them.

The Core Finding: Position Evaluation Fails, Pattern Analysis Works

Give an LLM a FEN string and ask “who is better and why”, and you get confident, fluent, frequently wrong answers. In my testing, GPT-5.2 agreed with Stockfish’s assessment (within ±0.5 pawns) on 61% of quiet middlegame positions. That number collapsed to 34% on tactical positions where a concrete sequence decided matters. Worse, the model almost never signalled its own uncertainty. A hallucinated knight fork is delivered in exactly the same tone as a correct one.

Here is a real example from one of my games. Position after 18…Rfe8, Black to move having just played it:

r2q1rk1/pp3ppp/2n1bn2/3p4/3P4/2N1BN2/PP2QPPP/2R2RK1 w - - 4 19

I asked GPT-5.2 to evaluate. It replied that White holds a pleasant edge of “around +0.6” because of the half-open c-file and pressure against c6, and suggested 19.Na4 heading for c5. Stockfish 17 gave 19.Na4 an evaluation of +0.18, essentially equal, and preferred 19.Rc2 followed by doubling. The narrative reasoning was fine. The number was invented. Note also that the position it described has a rook on f8 that the FEN does not contain in the way the model discussed it. That kind of quiet board-state drift happened in about one game in five.

Now flip the task. I fed the same model my last 30 losses as PGN with clock times attached, and asked it to find recurring themes rather than evaluate anything. The output was the most useful piece of chess feedback I have received in two years of play.

What Actually Worked: Move Times As The Input

The insight that changed my results: LLMs are bad at chess and good at data. Lichess exports clock times in the PGN by default. Nobody uses them. They are a direct readout of where your understanding runs out.

The prompt I settled on, which you can copy verbatim:

I’m going to paste 20 of my games in PGN with clock times. Do not evaluate any positions or suggest moves. Instead, for each game find the first move where I spent more than 90 seconds, and the first move where I spent under 5 seconds in a position with more than three plausible candidate moves. Tabulate: game number, move number, phase, my time spent, and the pawn structure or opening on the board. Then tell me what the table has in common.

The tabulated output across my 20 games looked like this (abbreviated):

GameFirst 90s+ movePhaseStructure
314early middlegameIQP, White side
512early middlegameIQP, White side
721middlegameCarlsbad
913early middlegameIQP, White side
1115early middlegamehanging pawns
1412early middlegameIQP, White side

Six of twenty long thinks landed on an isolated queen’s pawn position, always on moves 12 to 15, always as the attacker. The model then said something a human coach would say: I did not have a plan for the IQP middlegame, only for the IQP opening. My rushed moves clustered in rook endings, which meant I was playing them on autopilot at a level where autopilot loses half a point a game.

Stockfish will never tell you that. Stockfish evaluates the position in front of it and has no memory of the other nineteen games.

The Hybrid Workflow, Step By Step

Here is the loop I now run weekly. It takes about 40 minutes.

Step one. Run your games through Lichess’s own analysis (free, uses Stockfish 17 NNUE in the cloud). Export the analysed PGN. You now have engine evaluations embedded as comments, which matters enormously: you are giving the LLM ground truth instead of asking it to invent ground truth.

Step two. Paste that annotated PGN into ChatGPT with an explicit framing: “The engine evaluations in this PGN are correct. Do not second-guess them. Your job is to explain, in human terms, what misunderstanding would cause a 1500-rated player to make the move marked as an inaccuracy on move 19.”

That instruction does most of the work. When the model has the eval, it stops hallucinating the eval, and its explanations become genuinely good. My hit rate on “is this explanation sound” jumped from around 50% on raw FENs to roughly 85% on engine-annotated PGN, judged by whether I could verify the reasoning against a concrete line.

Step three. Ask for the anti-pattern, not the correction. Rather than “what should I have played”, try: “Give me three positions from master games where the same structural feature appears and the correct plan is clear. Name the games.”

This is where you need to be careful. Game citations hallucinate badly. Of 15 games GPT-5.2 named in my testing, 11 existed as described, 2 had the wrong year or wrong opponent, and 2 did not exist at all. Check every one against a database before you study it. Chessgames.com or the Lichess masters database takes ten seconds per game.

Step four. Turn the pattern into a drill. Ask the model to write you a study plan with a concrete stopping condition: “Twelve IQP positions from the Lichess masters database, White to move around move 14. I look at each for two minutes, write my plan, then check against the game continuation. Repeat until I’m right nine times out of twelve.”

The broader landscape of which models do what, and how the specialist tools compare, is covered in our pillar on AI coaching tools and LLM chess prompts, tested. This page is deliberately narrow: it is about whether ChatGPT specifically can do game analysis, and what that costs you in verification time.

Where It Fails, Precisely

Be clear about the failure modes so you can catch them live.

Illegal moves in suggested lines. In long variations, roughly one in four sequences of six-plus moves that I asked for contained at least one illegal or impossible move. The model loses track of the board. Short lines of two or three moves were far more reliable, maybe one error in fifteen.

Opening theory that sounds right. Ask for the main line of the Najdorf English Attack and you will get something that is 80% correct with a transposed move order or an interpolated move that quietly changes the position. If you memorise it, you will play it and be worse. Use an opening explorer for theory. Use the LLM to ask why the theory exists.

Endgame technique. Asked to convert rook and pawn versus rook, all three models I tested produced plausible-sounding prose about cutting the king off and building a bridge, and then gave move sequences that failed to actually win. Tablebase positions are solved. There is no reason to ask a language model.

Confidence calibration. This is the one that costs club players the most. The models rarely say “I’m not sure”. A prompt suffix helps: “Mark each claim as VERIFIED (you can state the concrete line), PATTERN (general principle, not calculated) or UNSURE.” Adding that suffix caused GPT-5.2 to flag about 30% of its own claims as PATTERN or UNSURE, and the flagged ones were where the errors lived.

A Worked Example End To End

Game 11 of my batch, a loss. Lichess flagged move 23 as a blunder, evaluation swinging from +0.4 to -2.1. Under the old workflow I would have seen the engine’s preferred move, nodded, and forgotten it by Thursday.

Instead I pasted the annotated PGN and asked: “I played 23.Bxf6 and the engine wants 23.Rfd1. Both keep material equal. What do I believe about bishops and knights that makes 23.Bxf6 look attractive to me, and why is that belief wrong in this specific structure?”

The reply, condensed: I was trading my good bishop to inflict doubled pawns, treating structural damage as the goal. In a closed position with pawns on both wings, the knight that remained was the better piece, and the doubled f-pawns actually opened the g-file for Black’s attack. The general rule (doubled pawns are weak) was overriding the specific assessment (which minor piece is better here).

That is correct, checkable, and it named a belief I actually hold. I then found four more games in my own history where I made the same trade. The pattern was worth about 60 rating points once I stopped.

Cost, Time And What To Expect

ChatGPT Plus at $20 a month plus free Lichess analysis covers everything described here. The paid Chess.com Game Review at $7 to $14 a month gives you engine analysis with canned prose explanations, which is faster but shallower: it explains moves, not you.

Budget 20 minutes per game for the full hybrid loop if you are doing it properly, including verifying cited games. That is expensive. Do it for two games a week, not twenty. The games worth it are the ones where you lost and do not know why, not the ones where you hung a piece on move 9.

One habit worth building: keep a running text file of every pattern the model identifies, and paste it back in at the start of each session. “Here are the eight weaknesses previously identified in my play. Does this week’s batch of games show improvement on any of them, and does it reveal a ninth?” Continuity is the thing most people never set up, and it is where the whole approach stops being a novelty and starts being a coach.