In August 2026 the full signal and pricing engine was cross-referenced against roughly 130 primary sources, from peer-reviewed journals to labeled practitioner research. The verdict set was published unedited: the core volatility-risk-premium discipline is the best-documented edge in the options literature, four subsystems were on the wrong side of the evidence, and the highest-priority fixes shipped days later in v3.96.0.
On August 16, 2026, the engine specification behind Options Scanner was cross-referenced, subsystem by subsystem, against roughly 130 primary sources: peer-reviewed journals first (Journal of Finance, Review of Financial Studies, Journal of Financial Economics, JFQA, Management Science), then working papers, then named practitioner research, which is labeled as such wherever academic coverage is absent. Every subsystem received a verdict from a fixed vocabulary: aligned, partial, deviation, divergent, or gap. Publishing the review is a bet on symmetry: the parts the evidence supports and the parts it did not are listed with equal prominence.
The engine's core discipline, selling defined-risk credit structures at 30 to 45 days to expiration on liquid underlyings when volatility is richly priced, is the single best-documented edge in the options literature: the volatility risk premium (Bakshi and Kapadia 2003; Coval and Shumway 2001; Carr and Wu 2009). Index put-writing benchmarks earned equity-like returns at roughly two thirds the volatility over 1986 to 2018, with an average implied-minus-realized gap of 4.2 vol points (Bondarenko 2019). Two honest caveats travel with that finding: the premium has compressed materially since 2006, and the evidence base is index and large-cap, which makes the engine's liquid-universe restriction load-bearing rather than cosmetic.
The review also rated as aligned: the execution stack (mid-anchored package limits, never legging into spreads, avoiding the open, and the natural price as the pessimistic bound: Muravyev and Pearson 2020; Chan, Chung and Johnson 1995), trusting the sign of dealer gamma exposure only at the index level (Barbon and Buraschi 2021), a tail posture of de-risking through sizing and halts rather than continuously held protective puts (Israelov 2019), fractional-Kelly sizing with a hard per-trade cap (MacLean, Thorp and Ziemba 2010), and the design of the ML validation stack (triple-barrier labels, purged cross-validation with embargo, deflated performance statistics: Lopez de Prado 2018).
The review named four divergences, published here in full:
Beyond the divergences, the review documented input gaps in the pricing stack: dividend yield fixed at zero (a 1 to 3 vol-point class error on quarterly payers near ex-dates), a hardcoded interest-rate input, no American-exercise treatment (early-exercise premia of 5 to 15% of value on in-the-money puts at 2026 rates: Barraclough and Whaley 2012), and calendar-day time handling that distorts the shortest expirations.
v3.96.0 shipped the wrong-side-of-evidence and input-correctness tiers within days of the review: NVRP, the engine's implied-minus-realized premium measure, replaced IV rank as the credit entry gate, fail-closed when the measure is unavailable; earnings calendars gained a hard gate requiring the implied announcement move to at least match the stock's historical announcement moves; a daily Treasury CMT curve replaced the hardcoded rate; dividend yield is now populated; an American-exercise overlay was added; delayed-feed pricing was refused on every traded path; and the probability-of-profit calibration report was promoted from diagnostic to a monitored gauge.
The review's remaining refinements are tracked rather than silently dropped, each with a reason: the reserved regime-model seam stays reserved until there is data to fit it honestly, heavier GARCH stacks were declined because the marginal forecasting value at this scale is negligible (Hansen and Lunde 2005), and the 50%-profit and 21-DTE exit conventions keep their honest framing as variance control rather than alpha, with the obligation to verify them on the engine's own fills.
The complete numbered bibliography, with all 130 sources, is in the full review: Engine vs. the Literature (PDF). Magnitudes above are quoted from the cited studies, gross of costs unless stated. Research commentary, not investment advice.