Methodology
How Prism Solver is validated
A solver is a measurement instrument: it is only worth using if it is calibrated. This page is the calibration record — what the engine computes, what it was checked against, and what the residual errors measure. Numbers are stated exactly as measured.
1 · Architecture
External-sampling MCCFR over card buckets
Prism Solver runs external-sampling Monte-Carlo CFR over a card-abstraction: hands are bucketed by river-equity percentile and, on earlier streets, by E[HS] / E[HS²] features over runouts. Card removal is exact — the sampler always deals real, non-conflicting hands from the actual ranges, so blocker effects during play are fully represented; only the strategy indexing is bucketed. Abstractions for postflop spots build in seconds at solve start and are disk-cached; bucket counts are configurable from a 2k/2k/2k laptop preset up to 16k/32k/64k on workstation RAM.
Why this design: multiway pot-limit games are far too large for the exact vectorized-CFR approach that works heads-up — the game tree and the per-node range matrices both explode with a third player and a four-card deal. Sampled CFR over abstraction buckets is the approach with a track record on exactly this problem class: it converges to an approximate equilibrium of the abstracted game, its error sources are measurable (sampling noise, under-training, abstraction coarseness), and each of those is quantified separately below rather than waved away.
2 · Engine certification
Certified against an exact solver
The sampled multiway engine is continuously validated against an exact heads-up oracle: a vectorized-CFR solver with no sampling and no card abstraction, maintained solely as a reference. On matched heads-up spots the two engines must agree, and do:
Game value agreement, sampled engine vs exact oracle: within 2% of pot.
On toy games small enough to enumerate, an exact best-response computation puts the sampled engine's exploitability below 1% of pot.
The distinction matters: the first check says the engine converges to the right value; the second says the strategy it converges to is actually hard to exploit, measured by full tree walk rather than by another sampled estimate. For production-size spots, where exact best response is infeasible, the engine reports a sampled best-response exploitability at end of solve.
3 · Tournament equity
ICM validated against exact enumeration
Tournament spots replace chip utilities with Malmuth-Harville equities. The textbook formula enumerates finish-order permutations, which is exponential in players; Prism computes it instead by an exponential-race integral, which scales to large-field stack models. The integral is validated against brute-force enumeration on stack configurations where enumeration is feasible:
Exponential-race ICM vs exact enumeration: agreement within 0.2% of the prize pool.
Bubble factors are derived from the same equities and printed with every ICM solve.
4 · Abstraction cost
The abstraction penalty, measured
Bucketing is the one approximation that never fully disappears with more training time, so we measured it directly. On a heads-up PLO river spot with uniform ranges we solved the same game twice: once with k = 2000 equity buckets (the default laptop preset) and once with no abstraction at all — all ≈178,000 distinct hands indexed individually. Both solves were run to convergence.
Self-play EV difference, k=2000 vs no abstraction, at convergence: 0.34% of pot.
The honest caveat: per-hand action frequencies differ more than EVs do (mean L1 distance ≈ 0.5 per hand). Abstraction preserves value while coarsening mixed-strategy detail — and since equilibria are not unique, two near-optimal strategies can legitimately mix the same hand differently at almost identical EV. A companion experiment (k=2000 vs k=8000 on a 3-way flop, deliberately stopped early at 3M iterations) showed much larger gaps — per-hand differences up to 0.48, seat EVs ±2.3% pot — dominated by under-training, not abstraction.
The practical guidance we ship with: quick solves are EV-indicative; for trustworthy frequencies, run until the stability metric drops below ≈ 0.01 (the app can pause itself there), and when a frequency drives a real decision, re-check it at a higher bucket count.
5 · External comparison
Head-to-head vs MonkerSolver in progress
Before public launch, Prism solves a fixed set of nine PLO spots — heads-up river, turn, and flop spots plus 3-way flops — that are also solved in MonkerSolver, the industry-reference multiway PLO solver, on byte-identical game trees. The protocol is designed so nothing is left to interpretation: uniform ranges to eliminate range-translation error, pot-fraction size menus that both tools express exactly, and raise-cap geometry chosen so both tools build the identical tree.
The acceptance gates, fixed in advance:
| Metric | River spots | Turn / flop spots |
|---|---|---|
| Per-seat EV difference | < 1% pot | < 2% pot |
| Query-hand EV difference | < 3% pot | < 5% pot |
| Root frequency L1 (aggregate) | ≤ 0.10 | ≤ 0.15 |
Results will be published here — all nine rows, including any misses, with commentary, both tools' versions and settings, the solve budgets used, and every spot configuration so the comparison can be reproduced. Any out-of-tolerance result blocks launch until it is explained.
6 · Reproducibility
Every number above is a test
The oracle agreement checks, the exact best-response bound, the ICM enumeration comparison, and the abstraction-cost measurement are not one-off experiments for this page — each ships as a reproducible test in the product's own suite and runs on every build. If a change to the engine moved one of these numbers, the build would fail before it reached you.