The Constant No Test Could Catch
A dispatch constant chooses which formula runs. Coverage says the line executed, mutation testing generates no mutants on it, and the suite passes either way. Only an independent reference can tell you the number is wrong — ours was, by 240x, at exactly the point it was switched on.
By MorphIQ Labs Research Notes
Every numerical library has them: a handful of bare floating-point constants that decide which formula runs. Below this value use the series expansion, above it use the closed form. They are usually inherited — from a paper, from a reference implementation, from whoever wrote the module first — and they are almost never verified, because there is no obvious way to verify them.
We held two of ours to an independent arbitrary-precision reference. One was wrong by a factor of 240 at precisely the point where it was switched on, and every test we had passed throughout.
Why a dispatch constant is unlike other code
Black-Scholes in its normalised form takes two arguments. With the forward, the strike, the time to expiry and the volatility, write
and then reduce to the pair that actually governs the numerics,
Here is moneyness measured in standard deviations and is half the total volatility. No single formula evaluates the option price well across the whole plane. Deep in the wings, where , the direct expression
subtracts two nearly equal quantities and loses most of its significant digits to cancellation. Near a different degeneracy appears. Production inverters therefore carry several evaluation paths and switch between them, and the switch points are bare constants.
What makes them different from ordinary code is that no test that compares the implementation to itself can adjudicate them.
- Coverage reports that the line executed. It cannot report that the branch it selected was the wrong one.
- Mutation testing generates no mutants on a
constdeclaration at all. We confirmed this directly: a mutation run over the module produced 1,220 mutants and not one of them perturbed a constant's value. - More fundamentally, mutation testing asks whether the suite would notice a change from current behaviour. When current behaviour is itself the defect, the question is circular.
- The existing suite asserted agreement to 4–10 ULP and passed for years, because its corpus contained no case near a branch boundary. A boundary is the only place a dispatch constant can be wrong.
The anchor
The only instrument that can settle it is a reference computed by a different route, at arbitrary precision, sharing no code with the implementation under test. We generate ours in a separate language from the engine, which makes accidental code sharing structurally impossible rather than merely discouraged.
Two disciplines matter more than they sound.
Arbitrary precision is not correctness. A cancelling formulation returns a badly wrong answer at any working precision — extra digits move the cliff without removing it. Our first oracle evaluated the normal CDF as
which returns exactly zero in the far left tail where the true value is around — silently, and precisely in the regime the oracle existed to certify. Rewritten with the complementary error function,
it recovers the tail exactly. We had already written a tolerance schedule from the broken version before catching it. The rule that came out of it: compute the same quantity two ways and require agreement at the extremes, not in the body.
An anchor can be too well-conditioned to measure anything. Construct inputs whose true answer is zero and see what the runtime reports — that is the noise floor below which no tolerance can distinguish signal. Build that experiment on a symmetric grid and the arithmetic comes out exact, the measured floor is identically zero, and the experiment has confirmed nothing while looking like a clean pass. Ask what would make a measurement come out clean for the wrong reason, and demonstrate it can fail before believing that it passed.
What the measurement showed
Sweeping the parameter each constant governs and comparing both candidate paths against the anchor produces the shape that gives the defect away — a cliff. Approaching the switch point from inside, the selected expansion degrades by orders of magnitude; past it, the displaced path takes over and error drops back to machine noise. A correctly placed constant sits where the two paths cross, with both at the noise floor. Ours sat well inside the decay.
The correction is a strict improvement in both directions, which is the uncomfortable part: for years the engine had been paying accuracy for nothing. The more accurate configuration also turned out to be 27% faster in the neighbouring case we examined, because the divergent expansion was carrying extra terms to chase an error it was creating.
Two outcomes, both good
We ran the same procedure against a second family of constants in the surface layer — tolerances deciding whether a fitted volatility surface is declared free of static arbitrage, a boolean that feeds an automatic repair pass. Those had no derivation recorded anywhere either.
They were correct. The measured noise floor sits many orders of magnitude below the tolerance in every realistic regime, and what the tolerance declines to flag is bounded far below anything economically meaningful. Nothing changed but the evidence.
This matters for the method's credibility rather than the engine's. A procedure that only ever finds defects is indistinguishable from motivated reasoning. The useful claim is not we find bugs — it is we can now tell the difference between a number that is right and a number that has merely never been questioned, and the second outcome is as much a result as the first.
What we changed permanently
The finding closed with four artifacts, and the fix itself is the least important of them.
- The derivation lives beside the constant — the measurement table, the criterion in words, and the headroom. A number without its criterion is not reproducible; "the largest value at which this expansion still beats the one it displaces" is.
- The anchor is committed and re-runnable, with its provenance recorded: authority, version, generation command, numeric type, conventions.
- A regression test was watched failing in both directions before being trusted. A gate never seen to fail is an intention, not a control.
- A findings register now records what went wrong and the measurement that proves it stopped. A closed entry names the measurement, never the commit: "fixed in !320" is an assertion, "72,245 ULP to 0.60 against an arbitrary-precision reference" is evidence.
The register did not exist before this finding. That is itself worth stating plainly — the defect was found from outside the repository, by an independently gated implementation holding this one to a reference it does not use, and there was nowhere in the engine to record it.
More from Ferro Suite
July 6, 2026 · 1 min
Wavelet Feature Extraction for Time-Series ML Pipelines
Time-series ML fails quietly when the feature layer ignores scale. A price series, volatility series, or sensor stream may look like one signal, but its behavior is usually a mixture of fast noise, medium-horizon stru…
July 6, 2026 · 2 min
What a Rust Options Pricing Engine Has to Prove
A Rust options pricing engine usually gets judged by surface area first: how many models, how many Greeks, how fast the implied-volatility solver runs. That is necessary, but it is not enough. The harder question is w…
July 6, 2026 · 1 min
Rust Market Data Normalization Starts at the Boundary
Market-data infrastructure becomes fragile when every downstream component learns a different venue dialect. One feed calls it a trade condition, another calls it a flag. One stream uses exchange timestamps, another c…