Validation
How we know the numbers are right. Every claim below names the check that enforces it; the full set runs before every release.
The discipline
The engines are held to bit-identical output: every change must reproduce the golden reference results exactly before it ships. A wrong number that looks right is the worst failure a numbers tool can have, so correctness is enforced by tests, not by review.
Cross-checks against published references
- Statistics: descriptive statistics, skewness and kurtosis verified against spreadsheet-standard formulas on classic textbook datasets; t-tests (pooled and Welch, including Welch's degrees-of-freedom rounding), two-sample z-tests and F-tests verified against published critical values (for example t(0.975,7) = 2.364624) and independent in-test computations that derive the answer a different way than the engine does.
- Distributions: every Monte Carlo sampler is checked against its theoretical mean and variance on 60,000 seeded draws, per distribution.
- Normal inverse: NORM.INV checked against jStat, an independent open-source statistics library.
- Optimization: linear programs with hand-verifiable optima (vertex enumeration written out in the test), two-phase minimization with inequality constraints, and nonlinear cases with closed-form solutions.
- Decision trees: expected-value rollback, optimal policy and risk profile verified against the classic oil-drilling textbook example.
- Project management: earned-value metrics (SV, SPI, CV, CPI and the verdict matrix) verified against a university course's worked example; critical-chain buffers on hand-computed networks.
- Schedule risk: the project simulation engine is pinned bit-for-bit against an independent reimplementation of the same network, so a change that moves any sampled duration or any percentile fails the build.
- Forecasting: the ETS and ARIMA forecasters carry pinned tests of their model selection, differencing choice and variance bands.
- Risk-aware optimization: every efficient-frontier point is checked bit-equal to the standalone optimization run at the same bound, under the same seed and common random numbers.
- Fourier analysis: FFT output checked against a known discrete Fourier transform and an independent reference implementation.
- Fuzzy matching: Jaro-Winkler similarity checked against published reference values.
Honesty is also tested
Marketing claims are guarded by the same tests: tests fail the build when a page claims something the code does not do, when a count on the site drifts from the registry, or when gate copy describes a limit that does not exist. The templates' worked notes quote figures computed by the engines themselves, and the tests recompute them.
Scale
As of 2026-08-21 more than 6,000 automated tests run on every commit. The engines run deterministically under a fixed seed, which is what makes the golden-reference discipline possible: the same model and seed always produce the same numbers, on any machine.