Engine Setup: GUIs, Settings And Cloud Depth
Most club players interact with Stockfish the way you’d interact with a vending machine. You press the button, a number falls out, you nod at the number, you move on. The number was produced at depth 18 with 16 MB of hash on a single core, it was the third-best move by a margin of 0.04 pawns, and you learned nothing. This page is about the other way to do it: running the engine as a tool you control, reading what it actually says, and pulling analysis from servers that have already spent a thousand times more compute on your position than your laptop ever will.
Everything here assumes Windows, because that’s where the binary-picking confusion lives. The UCI and GUI material applies equally on macOS and Linux.
How To Install Stockfish On PC Without Picking The Wrong Binary
The official download page at stockfishchess.org/download/ hands you a list of Windows builds with names like stockfish-windows-x86-64-avx2.exe, stockfish-windows-x86-64-bmi2.exe, and stockfish-windows-x86-64-vnni512.exe. These are the same engine compiled against different CPU instruction sets. Pick one your CPU doesn’t support and the program dies instantly with no useful message. Pick one that’s merely suboptimal and you quietly lose 20 to 40 percent of your node rate forever.
Find out what you have. In PowerShell:
Get-CimInstance Win32_Processor | Select-Object Name, NumberOfCores, NumberOfLogicalProcessors
That gives you the model string. Match it against these rules, which cover almost every machine a club player owns:
- Intel 4th gen (Haswell, 2013) through 10th gen: use
avx2. - Intel 11th gen and later desktop, or Ice Lake / Tiger Lake laptops: try
vnni512, fall back toavx2if it refuses to start. - AMD Ryzen 1000, 2000, 3000 (Zen 1 and Zen 2): use
avx2, notbmi2. These chips implement the PEXT instruction in microcode and it is dramatically slower than the software fallback. This is the single most common self-inflicted performance wound in engine setup. - AMD Ryzen 5000 (Zen 3):
bmi2is now fast, use it. - AMD Ryzen 7000 / 9000 (Zen 4, Zen 5):
vnni512. - Anything older than 2013, or a low-power Celeron/Atom:
sse41-popcnt, or plainx86-64if that fails.
Verify in one command. Open a terminal in the folder and run:
.\stockfish-windows-x86-64-avx2.exe bench
bench runs a fixed 50-position suite and prints a summary. A modern six-core desktop on a single thread will finish in roughly 2 to 4 seconds and report something in the region of 1.5 to 2.5 million nodes per second. If it crashes or exits silently, you chose too aggressively. Step down one level and try again. Compare the Nodes/second line between two candidate binaries and keep the faster one.
Three Windows-specific things that waste people’s evenings:
SmartScreen and Defender. The first launch may be blocked. Right-click the .exe, Properties, tick Unblock, Apply. Then add a Defender exclusion for the containing folder (Windows Security → Virus & threat protection → Manage settings → Exclusions → Add a folder). Without the exclusion, Defender re-scans the 70 MB binary on every process launch, and GUIs that restart the engine per position will feel sticky.
Folder location. Put it in C:\Engines\Stockfish\. Do not put it in Documents or Desktop if those are OneDrive-synced. Some GUIs write log and config files next to the binary, and sync conflicts on those files produce failures that look like engine bugs.
The neural net. Official release builds have the NNUE evaluation file embedded, so there is nothing else to download. If you ever compile from source or grab a development build, you may need to place a .nnue file alongside the binary and point EvalFile at it. A build that starts but evaluates the opening position at something absurd like +3.00 is a build that lost its net.
The Five Minutes In A Terminal That Teach You UCI
Before you attach the engine to any GUI, talk to it directly. Every GUI is a thin wrapper over this text protocol, and knowing the protocol means you can diagnose any GUI. Launch the executable from PowerShell and type uci:
Stockfish 17.1 by the Stockfish developers (see AUTHORS file)
uci
id name Stockfish 17.1
id author the Stockfish developers (see AUTHORS file)
option name Debug Log File type string default
option name NumaPolicy type string default auto
option name Threads type spin default 1 min 1 max 1024
option name Hash type spin default 16 min 1 max 33554432
option name MultiPV type spin default 1 min 1 max 500
option name Skill Level type spin default 20 min 0 max 20
option name UCI_LimitStrength type check default false
option name UCI_Elo type spin default 1320 min 1320 max 3190
option name UCI_ShowWDL type check default false
option name SyzygyPath type string default <empty>
uciok
Read the two defaults that matter: Threads 1 and Hash 16. Out of the box, on a machine with eight cores and 32 GB of RAM, Stockfish uses one core and sixteen megabytes. Most GUIs override this, several do it badly, and a few silently reset it when you change engines. Those two numbers deserve their own treatment, which is why they have their own page: Stockfish hash and threads settings that actually change your analysis covers how to size hash against your search length and why adding threads is not a linear speedup.
Now set things up and analyse a real position. This is the Two Knights Defence after 1.e4 e5 2.Nf3 Nc6 3.Bc4 Nf6 4.Ng5 d5 5.exd5 Na5:
setoption name Threads value 6
setoption name Hash value 2048
setoption name MultiPV value 3
setoption name UCI_ShowWDL value true
position startpos moves e2e4 e7e5 g1f3 b8c6 f1c4 g8f6 f3g5 d7d5 e4d5 c6a5
go depth 28
The engine streams info lines as it works, then stops:
info depth 28 seldepth 38 multipv 1 score cp 26 wdl 141 823 36 nodes 41238190
nps 5847312 hashfull 386 time 7053 pv c4b5 c7c6 d5c6 b7c6 b5e2 h7h6 g5f3 e5e4
f3e5 f8d6 d2d4 e4d3 e5d3 d8c7 h2h3 e8g8
info depth 28 seldepth 34 multipv 2 score cp 9 wdl 108 852 40 ... pv d2d3 h7h6
g5f3 e5e4 d3e4 a5c4 d1d4 ...
info depth 28 seldepth 33 multipv 3 score cp -4 wdl 92 848 60 ... pv c4f1 h7h6
g5f3 e5e4 d1e2 a5c4 ...
bestmove c4b5 ponder c7c6
Every field there is something you will misread at least once. depth 28 is the nominal search depth; seldepth 38 is how far the engine looked down the sharpest line. score cp 26 is +0.26 from the side to move. wdl 141 823 36 is win/draw/loss in per mille, also from the side to move: 14.1% White wins, 82.3% draws, 3.6% Black wins. hashfull 386 means the transposition table is 38.6% occupied, which is comfortable. nps 5847312 is your actual throughput, and it’s the only honest benchmark of your setup.
The other commands worth memorising: go movetime 10000 for a ten-second budget, go nodes 50000000 for a hardware-independent fixed amount of work, go infinite followed by stop for open-ended analysis, position fen <fen> to jump straight to a position, and ucinewgame to clear state between unrelated positions.
Choosing A GUI: What Each One Is Actually Good At
There is no single best GUI. There are four jobs (analysing a position deeply, reviewing a whole game, managing a database, playing sparring games) and different programs are good at different ones.
| GUI | Cost | Strongest at | The catch |
|---|---|---|---|
| Nibbler | Free | Live infinite analysis with MultiPV shown as a ranked, clickable list; WDL display built in | No database, no game review report, PGN handling is minimal |
| En Croissant | Free | All-round modern workflow: installs engines for you, generates accuracy reports, has an opening explorer and database | Younger project, occasional rough edges on very large PGN files |
| BanksiaGUI | Free | Running two or three engines side by side on the same position; engine-vs-engine matches with a readable UI | Interface density takes an hour to get used to |
| Arena 3.5.1 | Free | Rock-solid UCI compatibility, fine-grained engine config, tournaments | Interface is from 2010 and looks it |
| Cute Chess | Free | cutechess-cli for scripted match play and testing settings against each other | Command-line first, not an analysis tool |
| Scid vs PC | Free | Serious database work: huge PGN collections, position search, repertoire trees | Analysis panel is functional, not pleasant |
| ChessBase 18 / Fritz 20 | Paid | Database depth, Let’s Check cloud evals, opening repertoire tooling | Cost, and a learning curve measured in weeks |
| Lichess analysis board | Free | Zero setup, cloud evals for free, shareable study links | Browser WASM engine runs far slower than native and caps hash |
A practical pairing for a 1400-rated improver: Nibbler for deep single-position work, Lichess studies for storing and annotating what you find, and En Croissant or Lichess server analysis for the first pass over a finished game. That’s zero money and about fifteen minutes of setup.
One thing to understand about the browser engine: Lichess runs a WebAssembly build of Stockfish in your tab. It is the same search, but sandboxed, and it typically reaches a fraction of native node rate with a constrained hash size. For a quick check of whether a move hangs a piece, fine. For deciding whether a pawn sacrifice in your repertoire is sound, run it natively at depth 35 and let it think for two minutes.
The Settings That Matter Before You Touch Hash
MultiPV. Set this to 3 for analysis and leave it there. A single-line engine tells you what to play; a three-line engine tells you what your options cost. When the top three moves are +0.31, +0.28, and +0.24, the position is not “winning for White”, it’s balanced and flexible, and the engine’s preference is noise. When they are +0.30, -0.45, and -0.80, you have found a genuinely critical moment worth studying. MultiPV costs speed (three lines is roughly 30 to 50 percent slower to a given depth than one), which is a trade worth making for analysis and never for playing strength.
UCI_ShowWDL. Turn it on. Centipawns are a hard unit to calibrate. “82.3% draws” is not.
SyzygyPath. Endgame tablebases give perfect play for positions with few enough pieces. The 3-4-5 piece set is about 939 MB and is worth every megabyte if you study endgames: it turns “+1.4, probably winning” into “mate in 37, and here is the only move that keeps it”. The 6-piece set is roughly 150 GB, and unless you have a spare SSD and a genuine endgame obsession, skip it. Point SyzygyPath at the folder, for example C:\Engines\Syzygy\3-4-5, and confirm the engine picked them up by watching tbhits climb above zero in the info lines during an endgame.
UCI_LimitStrength and UCI_Elo. For sparring, set UCI_LimitStrength true and UCI_Elo to roughly 150 points above your own rating. The range is 1320 to 3190. A capped Stockfish plays a strange, lurching sort of chess, so treat it as a tactics-avoidance drill rather than a realistic opponent. Skill Level (0 to 20) is the older, blunter version of the same idea.
Move Overhead. Only relevant if the engine is playing timed games. 30 to 100 ms is sane; 10 ms will lose games on time to network or GUI lag.
Threads and Hash are the two with the biggest effect on what you actually see, and the two most often set wrong. The short version is that hash should be sized to the length of the search, not to your RAM, and that thread count interacts with how you read the eval. The long version, with measured numbers, is at /stockfish-hash-threads-settings/.
Reading The Evaluation Without Lying To Yourself
Modern Stockfish normalises its output so that +1.00 corresponds to roughly a 50% chance of winning the game from that position (calibrated around move 32 in a game between equally strong opponents). This is the fact that reframes everything. A +1.00 is not “a pawn up and cruising”. It is a coin flip that you happen to be favoured in. A +0.50 is a pleasant edge. A +2.00 is where “should win” starts.
Depth is not uniform currency. Depth 20 in a quiet manoeuvring middlegame is worth more than depth 20 in a sharp tactical mess where seldepth is 45 and every line ends in a queen sacrifice. Watch for the eval changing sign or jumping by more than 0.60 between consecutive depths. That instability is the engine telling you the position is genuinely unresolved, and it is a much stronger signal than the final number it happens to land on.
Three habits that separate useful analysis from vending-machine analysis:
Play out the principal variation on the board before you accept the eval. If the PV ends with a position you’d assess as clearly better for the other side, the engine has seen a resource you haven’t, and finding it is the actual training. Truncated PVs (a line that ends abruptly after four moves) usually mean a hash-table cutoff, not a short line.
Compare against the same nodes, not the same depth, when you’re testing anything. go nodes 20000000 on two different machines is a fair comparison. go depth 25 is not, because depth is influenced by hash size and thread count.
Never accept a single number on a critical position. Run it twice: once for ten seconds, once for two minutes. If the two agree, believe it. If the two-minute run reverses the ten-second run, the position belongs in your study file.
Cloud Depth: When Somebody Else Has Already Done The Work
Your laptop reaching depth 30 on a middlegame position in ninety seconds is respectable. A distributed database that has already analysed that exact position to depth 50 with billions of nodes is better, and it’s free.
Lichess cloud evaluations are served for any position that has been analysed enough times to be cached. On the analysis board they appear automatically as “Cloud analysis”. You can query them directly:
curl "https://lichess.org/api/cloud-eval?fen=rn1qkbnr/ppp2ppp/3p4/4p3/2B1P1b1/5N2/PPPP1PPP/RNBQK2R%20w%20KQkq%20-%200%205&multiPv=3"
The response is JSON with depth, knodes, and an array of pvs, each with a cp score and a move list. Positions in mainstream openings commonly come back at depth 40 to 50. A 404 means the position isn’t cached, which for anything past move 12 in an offbeat line is the normal answer.
chessdb.cn takes a different approach: a curated, human-and-engine-extended tree of evaluated positions. Query it with action=queryall and a board FEN and you get every known move from that position with a score attached. It is often deeper than Lichess cloud in opening theory, and it covers transpositions well.
Lichess server analysis (the “Request a computer analysis” button on a finished game) runs your whole game through volunteer-donated compute. It is tuned for throughput across thousands of games, not for depth on one position, so treat its output as a triage pass: it reliably finds the move where you dropped 2.5 pawns, and it will sometimes mislabel a subtle positional choice as an inaccuracy. Use it to locate the three moments worth studying, then analyse those three locally at depth 35.
ChessBase’s Let’s Check pools evaluations from paying users and is genuinely deep on opening positions, if you’re already in that ecosystem.
The dividing line is simple. Cloud wins for opening positions, common structures, and anything that thousands of other people have also looked at. Local wins for your own weird middlegame from last Tuesday’s league game, for tablebase endgames, and for anything where you want to control MultiPV and watch the eval evolve in real time.
A Worked Example: One Position, Four Ways
Back to the Two Knights. After 1.e4 e5 2.Nf3 Nc6 3.Bc4 Nf6 4.Ng5 d5 5.exd5, a 1200-rated player plays 5…Nxd5 and gets the Fried Liver: 6.Nxf7 Kxf7 7.Qf3+.
Vending-machine analysis: Lichess says the position after 7.Qf3+ is about +1.0 for White, so 5…Nxd5 was a blunder, don’t do that. Two seconds, nothing learned.
Interrogated analysis: run it locally with MultiPV 3 from the position after 5…Nxd5 and the picture changes. White’s +1.0 after the sacrifice, under normalised eval, means White wins roughly half the time. Black is not lost. Black is defending a grim but holdable position, and the practical question is whether you can hold it, which the number does not answer. Now compare it with 5…Na5, which came back at +0.26 with 82.3% draws in the output above. The difference between those two moves is not “bad versus good”. It is “one coin flip you’re behind on” versus “a position where four out of five games are drawn”. That’s a repertoire decision with a number attached to it.
Cross-checking against the cloud: query the post-7.Qf3+ position via the Lichess cloud API and you get a much deeper eval on the same position, plus the top three defensive tries ranked. If the cloud’s depth-48 eval broadly matches your local depth-30 eval, you can trust your local setup on positions like this. If it doesn’t, your hash is probably too small for the search length, and the hash and threads page explains why that specifically distorts long searches.
Sparring it: set UCI_LimitStrength true, UCI_Elo 1600, and play the Black side of the Fried Liver ten times against the engine. You will find out in about forty minutes whether “holdable” means holdable for you. That is the step almost nobody takes, and it is the one that converts an evaluation into a skill.
Turning Engine Output Into A Training Plan
Keep a blunder log. One file, one line per error, in this shape: position FEN, the move you played, the move the engine preferred, the eval swing in pawns, and one sentence in your own words about what you failed to see. Not “I missed Nxe5”, but “I assumed the knight was pinned and never checked whether the pin was real”. The second version is a pattern you can train; the first is trivia.
Filter ruthlessly by eval swing. Anything under 0.30 pawns is noise at club level and belongs in the bin. Between 0.30 and 1.00, you have a positional or planning error, which usually needs study rather than drilling. Over 1.00, you have a calculation or vision failure, which responds to tactics volume. Sorting your errors into those three buckets after twenty games will tell you where your next three months should go, and the answer is frequently not the one you expected.
Run the guess-the-move drill on master games. Load a game in a GUI that supports engine analysis, hide the moves, and at each position write down your candidate before advancing. Use MultiPV 3 to check: if your move is in the top three, you are fine even when it isn’t first. Track the percentage over fifty positions. A 1400 player typically lands in the top three around 45 to 55 percent of the time in quiet positions and far less in sharp ones, and watching that number climb is a better progress signal than rating over any period shorter than six months.
Use the cloud for repertoire work and the local engine for post-mortems. Building your Caro-Kann files against depth-45 cloud evals costs nothing and takes minutes. Understanding why you lost the rook endgame on Tuesday requires your own machine, tablebases loaded, and twenty minutes of walking the PV backwards until you find the move where the draw evaporated.
Set a recurring calendar block, thirty minutes, twice a week, labelled with the specific thing you’re checking. “Engine work” is not a plan. “Review the four positions from the blunder log tagged missed opponent’s counterplay, at depth 35, MultiPV 3” is a plan, and you’ll actually do it.
In this section
The supporting pages under this subject.