Merged master report — synthesized from three
independently-commissioned deep-research reports (Qwen, Gemini,
MiniMax), commissioned against the CASINO prompt at
casino-prompt-100-to-200.md. Research date: 2026-08-01. A
fourth supplied report (Perplexity) was excluded as a content-free stub
— see Section 1.
Verification status: this report resolved approximately 230 cross-report conflicts during merge, including one instance of a fabricated citation (struck — see Section 5), and downgrades every claim it could not verify against a specific, citable source rather than presenting it as settled. Table-level and citation-level detail was extracted from the three source reports and their structured digests; regulatory, tax, and package-recency claims are carried forward from the source reports' own stated verification rather than independently re-queried against primary registries during this merge pass — see Section 1 and Section 13 for exactly which claims that caveat applies to.
Research date: 2026-08-01. Commissioning
prompt: the CASINO deep-research prompt at repository root
casino-prompt-100-to-200.md, targeting a USD 100 to USD 200
growth objective within approximately 90 days for a Massachusetts retail
investor.
Source reports merged: three independently-commissioned deep-research reports, generated by three different AI research tools against the identical commissioning prompt — referred to throughout this report as Qwen, Gemini, and MiniMax. A fourth report (Perplexity) was supplied but contained no substantive content — a title, one reference, and a literal placeholder string in place of a body — and was excluded from the merge entirely rather than counted as a corroborating or dissenting source.
Merge process: each source report was digested
individually, then merged section-by-section against the original
prompt's 10 Intent clusters using majority-rule conflict resolution
(2-of-3 agreement wins; where no majority existed, either both versions
were dropped as irreconcilable or a range/qualified statement was
retained; citation quality broke ties). Approximately 230 individual
conflicts were identified and adjudicated across the 10 content
sections; full detail is in
output/notes/phase3-conflict-resolution.md and in each
section's own "Resolved Conflicts" subsection. One notable outcome of
this process: a set of Polymarket favorite-longshot-bias figures in the
MiniMax report were traced to a fabricated citation and struck from the
merged findings (see Section 5's Resolved Conflicts and the top-level
output/conflict-resolution.md).
Sources consulted: the merged bibliography (Section 14) indexes approximately 105 distinct citations — roughly 86 DOI-bearing academic and empirical sources, 13 regulatory/statutory primary sources, and the remainder non-DOI academic sources (books, working papers, preprints) — drawn from across the three source reports' own citation bases. This report does not independently browse or re-verify primary sources beyond what the three source reports themselves cited; where a source report's own citation-integrity self-check flagged a claim as unverifiable or fabricated, that flag was preserved and the claim was downgraded or struck during merge rather than passed through as corroborated.
Evidence-tier distribution: claims are tagged inline
throughout Sections 3–12 using the T1–T6 taxonomy defined in Section
14's introduction. The two densest sections in raw citation count —
Section 5 (strategies with peer-reviewed support) and Section 6
(outside-finance forecasting techniques) — carry the largest share of
T1/T2 tags; Section 11 (regulatory/tax) carries the largest share of
claims deliberately downgraded to "unsettled"/T6 out of legal caution
rather than asserted as settled on thin sourcing. No exhaustive
tier-count tally across all 10 sections was performed as part of this
merge; a reader requiring an exact count should grep the assembled
report for [T1] through [T6] tags
directly.
Questions from the original prompt's Intent section that could not be answered, and why:
median_days_to_target field in the JSON appendix is
therefore null rather than a fabricated number. This is
flagged as the single weakest point in the merged report by its own
adversarial self-critique (Section "How This Report Could Be
Wrong").python_libraries array.No candidate strategy in the surveyed universe carries positive expected value after trading costs and taxes at USD 100 scale. This is a unanimous finding across all three independently-commissioned source reports underlying this merge — not a majority-rule adjudication, but genuine convergence. The only positive-expected-value result anywhere in the merged evidence is a gross, pre-cost, pre-tax figure on long-tail event-contract longshots, and it is explicitly too small in magnitude to plausibly compound to a 100% return within a 90-day window even before costs are subtracted.
The central estimate for reaching USD 200 from a USD 100 stake within 90 days, under the best-supported approach identified (unlevered spot cryptocurrency exposure — ranked highest not because of any documented statistical edge, but because it is the only unlevered, friction-cheap vehicle with genuinely uncapped upside), is approximately 3%, within a defensible band of 1%–8%. The corresponding aggregate probability of ruin — defined here specifically as terminal wealth falling to USD 25 or below, a threshold chosen because it represents an experiment-ending loss rather than literal zero — is approximately 60%, within a band of 45%–75%. Literal total loss (terminal wealth of exactly USD 0) is near zero for unlevered spot positions specifically, but reaches 55%–92% for long-premium options strategies and 70%–90% for event-contract longshot bets — the vehicle matters enormously to which ruin figure applies.
Where the evidence is thinnest, and what that means for trusting these numbers: none of the three source reports derives an actual first-passage time distribution for any individual strategy — all three report terminal-return distributions only, meaning every specific "P(reach $200) = X%" figure in this report is model output stapled onto a return estimate that was never designed to answer the time-bounded question actually being asked. This is documented in this report's own adversarial self-critique as its single weakest point. Treat the probability figures above as directionally reliable — the ordering (crypto and equities feasible-but-low-odds; leveraged and short-premium vehicles worse; martingale sizing and technical-analysis-driven day trading catastrophically worse) is robust across all three source reports — but do not treat the specific percentages as precision estimates.
What does not work, documented with the same rigor as what might: retail day trading (fewer than 1% of participants show persistent, predictable profitability net of fees, per the Taiwan and Brazil day-trader literatures); technical-analysis pattern trading (a multiple-testing-corrected study of over 15,000 trading rules found zero survive out-of-sample after costs); leveraged and inverse ETFs held beyond a single day (volatility-decay compounding that both source reports needed correcting during merge, since the original MiniMax formula was roughly double the correct magnitude); penny stocks and OTC securities (extreme spreads and pump-and-dump structuring); social-media and meme-momentum signals; naive machine learning applied to price series without purged cross-validation; copy-trading and paid signal services; and martingale-style progressive position sizing, which drives probability of ruin toward certainty over a long enough sequence given a finite USD 100 bankroll. This report's "strategies that do not work" section (Section 7) is, by design and by verified word count, not shorter than its "strategies with peer-reviewed support" counterpart (Section 5) — negative findings received equal editorial weight, not an afterthought treatment.
A fabricated citation was caught and removed during this merge, not before. One source report's favorite-longshot-bias claims for Polymarket rested entirely on a citation to a paper that does not exist, attributed to an author whose surname matches this project's commissioner — evidently pattern-matched from context rather than sourced from anything real. All findings resting on that citation were struck rather than passed through as corroborated by majority rule (see Section 5's Resolved Conflicts). This report's own adversarial self-critique treats this as a structural warning: majority-rule reconciliation across three AI-generated research reports does not protect against correlated fabrication, since all three tools may share similar failure modes, and in this instance only single-source isolation — not agreement — caught the problem.
Vehicle and regulatory feasibility: fractional equities/ETFs, listed single-leg options, Kalshi and ForecastEx event contracts, and spot cryptocurrency are all technically executable at USD 100 scale with sub-10% round-trip friction under the right order-routing choices (notably: use advanced/pro crypto order interfaces, not "simple trade" retail interfaces, to avoid a roughly 4x fee markup). The Massachusetts legal status of CFTC-regulated event contracts is the most consequential open regulatory question in this report: a state-court preliminary injunction against a major event-contract venue was pending State Judicial Court review as of the research date, in direct tension with the CFTC's asserted federal preemption — no source report could supply a definitive, currently-resolved answer. A Massachusetts resident should not assume uncontested lawful access to non-sports event contracts pending that resolution. Federal and Massachusetts tax characterization of prediction-market proceeds specifically remains genuinely unsettled in primary law, not merely under-researched, and the two candidate characterizations (capital gain versus gambling winnings) diverge sharply on whether losses are deductible at the Massachusetts state level.
Software and data infrastructure: the merged Python
stack (Section 9, Table E) spans data ingestion, time-series/econometric
modeling, Bayesian inference, forecasting, backtesting, portfolio/Kelly
sizing, options pricing, and finance-appropriate (purged/combinatorial)
cross-validation, all built from packages the source reports report as
actively maintained within the trailing 365 days — several widely-cited
alternatives (backtrader,
zipline/zipline-reloaded,
pyalgotrade, pyfolio, opstrat,
properscoring) are explicitly flagged as abandoned and
excluded from the recommendation set. Free data sources (Section 10,
Table F) cover equities, macro series, SEC filings, and event-contract
order books, but the free tier of the most commonly used equity-data
source (yfinance) is neither point-in-time nor
survivorship-bias-free, a defect Section 8's backtesting-integrity
discussion identifies as one of the most common ways a retail backtest
silently produces an inflated, non-reproducible result.
Bottom line for the decision this report exists to inform: the evidence does not support an expectation of doubling USD 100 within 90 days through any legal vehicle surveyed. If the experiment proceeds regardless — as an explicitly bounded, fully-loss-tolerant exercise rather than an investment expected to succeed — Section 12's epistemic salvage plan specifies what to measure so the 90 days produce a generalizable lesson even when, as expected, the capital objective is not met.
All three source reports converge on the same opening move, and it is the correct one: "turn USD 100 into USD 200 within 90 days" is not a rate-of-return question. It is a first-passage problem — the probability that a stochastic wealth process touches an upper absorbing barrier before it touches a lower one, inside a hard terminal time.
This distinction is not cosmetic. CAGR, the Sharpe ratio, arithmetic
expected return, and Jensen's alpha are all functionals of the
central tendency and dispersion of the terminal wealth
distribution. The objective here is a functional of the upper tail's
hitting mass under a deadline. Two portfolios can carry identical
Sharpe ratios and differ by an order of magnitude in P(reach 2×).
Optimizing the first tells you almost nothing about the second. Any
strategy comparison presented on a CAGR or Sharpe axis is answering a
different question than the one posed. [T6] (Source:
Qwen, Gemini, MiniMax — unanimous in direction; the specific worked
counterexample MiniMax offers to illustrate it is arithmetically false
and is excluded — see Resolved Conflicts C-17.)
(Source: Qwen, Gemini, MiniMax)
Let W_t denote the wealth process,
W_0 = 100, with an absorbing target barrier at
W = 200 and an absorbing ruin barrier at
W = 0. Define the hitting times
τ₊ = inf{t ≥ 0 : W_t ≥ 200},τ₀ = inf{t ≥ 0 : W_t ≤ 0}
and the deadline T = 90 calendar days ≈ 63 US trading
days ≈ 13 weekly intervals. The decision is governed by exactly three
scalars:
P₊ = P(τ₊ < min(τ₀, T))— successP₀ = P(τ₀ < min(τ₊, T))— ruinP_T = 1 − P₊ − P₀— neither barrier reached by the deadline
plus the conditional distribution of τ₊ | (τ₊ < T) —
the time-to-target law. All three reports supply the stopping-time
construction; Qwen and MiniMax state it in wealth space, Gemini in log
space. [T1] on the mathematics — this is the standard
two-barrier gambler's-ruin construction in Feller, An Introduction
to Probability Theory and Its Applications, vol. 1, ch. XIV, as
named by MiniMax (book, no DOI). [T6] on the sourcing: Qwen
and Gemini attach no citation, and Feller never reaches MiniMax's own
source list. (Source: Qwen, Gemini, MiniMax)
Under linear utility over the terminal stake — appropriate only where
the principal has genuinely pre-committed to absorbing the full USD 100
loss without material consequence — the value is
V = 200·P₊ + 0·P₀ + 100·P_T, which reduces to
V − 100 = 100·(P₊ − P₀). That algebra is correct and is
worth stating because it makes the symmetry explicit: under this
utility, a strategy is worth pursuing only insofar as it moves
P₊ faster than it moves P₀. [T6]
(Source: MiniMax only)
The log-space diffusion formulation. Write
X_t = ln(W_t/W_0) under arithmetic Brownian motion with
drift θ = μ − ½σ² and diffusion σ:
dX_t = θ dt + σ dW_t,X_0 = 0
The target boundary sits at
B = ln(200/100) = ln 2 ≈ 0.6931; the ruin boundary at
A = ln(ε/100) → −∞. Then:
P(τ_B < τ_A) = [1 − e^(2θA/σ²)] / [e^(−2θB/σ²) − e^(2θA/σ²)]A → −∞ with θ ≤ 0:
P(τ_B < ∞) = e^(2θB/σ²) = 2^((2μ/σ²) − 1) ≤ 1f(t) = B/√(2πσ²t³) · exp(−(B − θt)²/(2σ²t))P(τ_B ≤ T) = Φ((θT − B)/(σ√T)) + e^(2θB/σ²) · Φ((−θT − B)/(σ√T))Gemini and MiniMax state the first-passage density in identical
functional form, independently — the strongest cross-report agreement in
this cluster. These are standard results in Karlin & Taylor, A
First Course in Stochastic Processes (1975), and Feller, An
Introduction to Probability Theory and Its Applications, vol. 1,
ch. XIV, as named by MiniMax. [T1] on the mathematics;
[T6] on the sourcing — Gemini attaches no citation
at all to any formula in this block despite tagging them
[T1], and neither Karlin & Taylor nor Feller reaches
MiniMax's own source list. Qwen declines the closed form entirely,
asserting only that the general case "requires numerical methods like
Monte Carlo simulation" — true for path-dependent or
non-constant-parameter processes, but it forgoes the analytics available
under constant drift and diffusion. [T4] (Source:
Gemini and MiniMax on the formulas; Qwen on the Monte Carlo
fallback)
The finite-horizon problem is a stochastic control problem,
not a barrier problem. Gemini alone states the correct general
formulation: maximize
P(τ_target ≤ T ∧ τ_target < τ_ruin) over admissible
strategies π, with the value function satisfying the
Hamilton–Jacobi–Bellman PDE
∂V/∂t + max_π{ μ(w,π)·∂V/∂w + ½σ²(w,π)·∂²V/∂w² } = 0
subject to V(w,T) = 1{w ≥ 200},
V(200,t) = 1, V(0,t) = 0. This matters for
§3.4 below. [T6] — Gemini supplies no citation, and its own
digest confirms none is attached. (Source: Gemini only)
What no source supplies. None of the three reports
produces a defensible numerical P₊, P₀, or
median τ₊ for this objective. MiniMax tabulates eight
parameterized return processes with point estimates and interquartile
ranges, but every row fails on inspection: two rows describe a ±50%
wager that can reach neither barrier, two carry an annualized volatility
mislabeled as daily, one reports a four-fold range in a column of point
estimates, the equity row is off by roughly six orders of magnitude with
the drift sign inverted, one applies the secretary problem outside its
domain, and one requires 160% leverage under a stated no-margin
constraint. Gemini's headline figures (P₊ = 0.18,
P₀ = 0.82) are stated to two significant figures with no
drift parameter, no volatility parameter, no trial count, no code, and
no citation. Qwen's ranking-table probabilities are round assertions
with no model, despite the report itself prescribing Monte Carlo.
The literature surveyed here does not contain a validated
first-passage probability for a USD 100 → USD 200 / 90-day objective,
and neither do these three reports. [T6]
(Source: all three, by exclusion)
The Kelly criterion for a binary bet with win probability
p, loss probability q = 1 − p, and net odds
b is
f* = (bp − q) / b
which at even money (b = 1) reduces to
f* = 2p − 1. The continuous-return analogue is
f* = (μ − r)/σ². [T1] — Kelly, J. L. (1956),
"A New Interpretation of Information Rate," Bell System Technical
Journal 35(4), 917–926, DOI 10.1002/j.1538-7305.1956.tb03809.x.
(Source: Qwen and Gemini agree on this form; MiniMax's variant is
arithmetically wrong and is excluded — see Resolved Conflicts
C-1.)
Kelly maximizes E[log W] — the asymptotic geometric
growth rate. Breiman's theorem (1961) establishes that log-wealth
maximization asymptotically minimizes the expected time to reach an
arbitrarily large wealth target as T → ∞. [T6]
on the citation: Gemini names Breiman's theorem, tags it
[T1], makes it load-bearing for its entire Kelly-inadequacy
argument, and then omits it from its own 47-entry bibliography. No
source in this merge supplies a venue, page range, or DOI for it, and
none is invented here.
All three reports agree Kelly is the wrong objective function for a fixed multiple under a deadline. The reasons, merged:
E[log W_T] = ∫ log w · dF_T(w). The first-passage objective
is E[1{τ₊ < T}]. These are different functionals of the
same wealth distribution and they produce different optima. Kelly is
silent on finite-horizon hitting probability: a strategy with a
small positive geometric growth rate can have arbitrarily small
probability of touching 2× inside 90 days. [T4]
(Source: Qwen, Gemini, MiniMax)ln 0 = −∞), which forces
conservatism structurally. The stated utility here is binary —
U(W_T) = 1{W_T ≥ 200} — and assigns zero marginal value to
the retained stake. A utility that is indifferent between USD 100 and
USD 0 cannot be optimized by a rule built to avoid USD 0 at any cost.
[T4] (Source: Gemini only — the sharpest single
formulation of the mismatch in any of the three reports)p ≈ 0.52) it fails to push probability
mass across W = 200 inside the horizon at all.
[T6] — Gemini states this without citation; it is a widely
repeated Kelly property in practitioner literature but no source here
sources it.μ
or σ² produces a strategy that sits closer to full Kelly at
the true parameters than intended, and estimation error is the
dominant source of long-run Kelly underperformance. [T4]
(Source: MiniMax, citing the MacLean–Thorp–Ziemba literature without
a specific work)Two quantitative claims about fractional Kelly must be quarantined.
MiniMax states that κ-fractional Kelly achieves
(1 − κ²) of full-Kelly log growth; that expression yields
zero growth at κ = 1 and is wrong. The
correct continuous-approximation factor is
κ(2 − κ) = 1 − (1 − κ)², which equals 1 at full Kelly; both
expressions coincidentally return 0.75 at half Kelly, which plausibly
masked the error. [T6] — corrected here by inspection,
uncited, and nothing in this section is built on it. MiniMax's further
claim that "a 20% relative misestimation of μ or
σ² is enough to make full Kelly worse than half Kelly for
any plausible parameter vector" carries a universal
quantifier on a specific threshold and is sourced only to "the
MacLean–Ziemba literature" with no named work. [T6],
unverified. (Source: MiniMax only, both claims)
One minority claim is pruned. Qwen asserts that Kelly "would likely lead to premature ruin by taking overly aggressive positions relative to the 90-day deadline." This is asserted, not derived; it contradicts Qwen's own subsequent endorsement of bold play; and it is analytically backwards for the environment Qwen itself posits — for a subfair game the Kelly fraction is zero or negative (do not bet), not aggressive. Excluded. (Source: Qwen only; contradicted by Gemini and MiniMax)
The defensible synthesis: Kelly is a ceiling on stake size,
not a selector of strategy. Choose the strategy class under the
hitting-probability criterion; size within it under fractional Kelly to
absorb parameter uncertainty. Do not invert that order.
[T4] (Source: MiniMax; consistent with Gemini)
The theorem. Dubins, L. E., & Savage, L. J.
(1965), How To Gamble If You Must: Inequalities for Stochastic
Processes, McGraw-Hill; ISBN 978-0486780641 (Dover reprint). Book,
no DOI. [T1] For the red-and-black subfair game — a random
walk on [0, M] with absorbing barriers at 0 and
M, win probability p < ½ — bold
play maximizes the probability of reaching M before
0, among all measurable stake-selection strategies. Bold play
is defined as
y_bold(x) = min(x, M − x)
that is: stake exactly enough to reach the goal, never overshoot. Timid play — repeated minimum-size wagers — gives the classical two-barrier gambler's-ruin value
P_timid = [1 − (q/p)^{W₀}] / [1 − (q/p)^{M}] ≪ p = P_bold
[T1] on the theorem. (Source: Qwen and Gemini state
bold play correctly as min(x, M − x); Gemini alone supplies
the timid-play comparison. MiniMax's characterization — "wagering the
maximum possible stake on every round" — and its gambler's-ruin formula
are both wrong and are excluded; see Resolved Conflicts C-2 and
C-3.)
At this specific starting point, the two readings
collapse. With W₀ = 100 and M = 200,
min(x, M − x) = min(100, 100) = 100. Bold play here
is the entire stake on a single trial, and it resolves in one
period, giving P_bold = p exactly. This coincidence holds
only because the principal starts at precisely half the target; it would
not survive a win. Worth stating because it is why MiniMax's
mischaracterization produces the right operational number at
t = 0 despite being the wrong rule. [T6] —
this reconciliation is the merge's own, not any source's.
The contradiction with diversification is real, and both
sides are correct. All three reports state it; Gemini names it
the "Diversification Paradox." The mechanism: under a fixed deadline in
a subfair or marginally-fair environment, E[W_T] ≤ W_0, so
the mean of the terminal distribution lies below the target.
Variance is then the only mechanism that supplies probability
mass above W = 200. Diversification dampens
variance, pulling the trajectory toward a mean that is below the target,
and thereby reduces P₊. This inverts the central
prescription of Modern Portfolio Theory. [T4] (Source:
Qwen, Gemini, MiniMax — unanimous)
Both prescriptions are correct because they optimize different functionals:
| Objective | Functional | Optimum in a subfair game |
|---|---|---|
| Reach a fixed multiple before a deadline | P(τ₊ < min(τ₀, T)) |
Concentrate — bold play |
| Maximize risk-adjusted terminal wealth | E[W_T]/σ(W_T), or E[log W_T] |
Diversify — Kelly-sized, spread across independent edges |
[T1] on both rows — Dubins & Savage (1965), How
To Gamble If You Must, ISBN 978-0486780641, for row 1; Kelly
(1956), DOI 10.1002/j.1538-7305.1956.tb03809.x, for row 2.
Neither is "the" right answer; each is the right answer to its own
question. The critical asymmetry the Narrator must state plainly:
the P₊-maximizing strategy is simultaneously the
P₀-maximizing strategy. Bold play is optimal only
relative to an objective that assigns zero value to the retained stake.
Under V = 100·(P₊ − P₀), the linear-utility value of a
single fair-coin bold bet is exactly zero, and of a subfair one,
negative. Concentration does not create expected value; it relocates
existing expected value from the middle of the distribution into the two
tails. [T6] — merge's own framing.
The superfair inversion. For p > ½,
and for the growth-rate objective, the optimum inverts:
fractional-Kelly sizing with diversification across independent edges
dominates. Qwen and MiniMax both state this. MiniMax adds the important
qualifier that the inversion applies to the growth objective and
not to the hitting objective — even a positive edge
does not make timid play optimal for P₊. [T4]
(Source: Qwen, MiniMax)
The empirical routing is what makes this matter. The
Barber–Odean, Barber–Lee–Liu–Odean, and Chague–De-Losso–Giovannetti
studies (§3.6) all find negative realized edge for the average
retail participant net of costs. The subfair case is therefore the
applicable one for an unspecialized retail principal, which is precisely
the regime where bold play maximizes P₊ and diversification
reduces it. [T1] — Barber, B. M., Lee, Y.-T., Liu, Y.-J.,
& Odean, T. (2014), "The Cross-Section of Speculator Skill: Evidence
from Day Trading," Journal of Financial Markets 18, 1–24, DOI
10.1016/j.finmar.2013.05.006; Barber, B. M., & Odean, T. (2000),
"Trading Is Hazardous to Your Wealth," Journal of Finance
55(2), 773–806 (DOI [T6], see §3.6); Chague, F., &
Giovannetti, B. (2025), "The COVID-19 and Day-Trade Pandemics in
Brazil," Brazilian Review of Finance 23(1), DOI
10.12660/rbfin.v23n1.2025.94291 [T2]. (Source: MiniMax,
with the empirical anchors; Qwen reaches the same routing without
figures)
A limitation none of the three sources states, and the merge
must. [T6] Dubins–Savage optimality is a result
about the unbounded-time goal problem: reach M
before 0, with no terminal date. This objective has a hard
deadline at T = 90, and the finite-horizon optimum
is the solution to Gemini's HJB equation with terminal condition
V(w,T) = 1{w ≥ 200} — which is not, in general, bold play.
MiniMax attempts to bridge this by asserting that a deadline
"effectively" introduces a time discount and then invoking Chen (1977)
to argue the bold-play optimum can invert. That bridge does not hold: a
geometric discount factor and a hard terminal time are different
modifications to the objective functional, and the equivalence is
asserted rather than shown. No source in this merge, and no
literature identified by any of the three, solves the bold-versus-timid
question under a hard finite horizon at these parameters. The
bold-play result should therefore be read as qualitative guidance about
the direction concentration pushes P₊, not as a
sizing prescription.
Chen's discount-factor inversion, which qualifies the
prescription further. Chen, R. (1977), "Subfair primitive
casino with a discount factor," Zeitschrift für
Wahrscheinlichkeitstheorie und Verwandte Gebiete 39, 167–174, DOI
10.1007/BF00535184. [T2] on the citation — the paper exists
and its topic matches. [T6] on the directional claim:
MiniMax reports that adding a time discount to a subfair game can make
timid play outperform bold play, uses that conclusion twice as the hinge
that walks back its own bold-play prescription, and its own
open-questions section concedes the quantitative implications for a
90-day horizon with realistic frictions "are not characterised in the
literature." The specific directional conclusion is unverified in this
merge and should be treated as such. (Source: MiniMax only)
Cramér–Lundberg. The classical surplus process is
U(t) = u + c·t − Σ_{i=1}^{N(t)} X_i
with initial surplus u, deterministic premium inflow at
rate c, claim arrivals N(t) ~ Poisson(λ), and
i.i.d. claim severities X_i. The ultimate ruin probability
is ψ(u) = P(inf_{t ≥ 0} U(t) < 0). Lundberg's
inequality bounds it:
ψ(u) ≤ e^{−Ru}
where the adjustment coefficient R is the unique
positive root of λ + cR = λ·M_X(R).
[T1] on the mathematics — standard results originating with
Lundberg (1903) and Cramér, with the canonical modern treatment being
Asmussen, S., & Albrecher, H. (2010), Ruin Probabilities,
2nd ed., World Scientific (book, no DOI), as named by MiniMax.
[T6] on the sourcing: Gemini attaches no citation to the
inequality, and Asmussen & Albrecher never reaches MiniMax's own
source list. (Source: all three state the surplus process; Gemini
alone states Lundberg's inequality with the correct root
condition.)
Under exponentially distributed claim severities with mean
μ, the closed form is
ψ(u) = (λμ/c)·exp(−(1/μ − λ/c)·u) when
c > λμ — the safety-loading condition.
At zero safety loading (c = λμ), ψ(u) = 1:
ruin is certain. That last result is the actuarial statement of the same
thing Qwen derives from gambler's ruin — in a game with no positive
expectation, ruin is not a risk, it is an eventual certainty; only the
timing is stochastic. [T1] on the mathematics (Asmussen
& Albrecher 2010, as above); [T6] on the sourcing —
MiniMax states the closed form with no citation attached at the point of
claim. (Source: MiniMax on the exponential closed form; Qwen on the
gambler's-ruin analogue)
The transfer to an investment process is analogical, not
rigorous, and all three sources overstate it. Cramér–Lundberg
is a one-sided model: deterministic upward drift punctuated by
compound-Poisson downward jumps. The investor's problem is two-sided and
diffusive, with no premium stream and with an upper absorbing barrier
the actuarial model does not have. All three reports gesture at the
mapping — "replace claims with losses and premiums with gains" — and
none does the work. The two-boundary generalization requires the Sparre
Andersen model and the Gerber–Shiu expected discounted penalty function
(Gerber, H. U., & Shiu, E. S. W. (1998), "On the Time Value of
Ruin," North American Actuarial Journal 2(1), 48–72, DOI
10.1080/10920277.1998.10595671, [T1]), which both Gemini
and MiniMax name and neither applies. [T6] on the
transfer's validity at these parameters. (Source: Qwen, Gemini,
MiniMax on the model; merge's own on the limitation)
Two MiniMax constructions in this area are excluded: its repurposing
of the Lundberg bound as an upper bound on P(reach H) — the
classical inequality bounds ruin, and the reversal is asserted
with no derivation — and its tabulated negative adjustment
coefficients, which are impossible under its own stated definition of
R as the unique positive root and which would make
e^{−R(H−x)} > 1, not a probability bound at all.
(See Resolved Conflicts C-9, C-10.)
Optimal stopping. Qwen and Gemini reach the same
conclusion by different routes, and merged it is a single point:
under a hard deadline with two absorbing barriers, the stopping
rule is degenerate. Formally, the Snell envelope
V_t = ess sup_τ E[Y_τ | F_t] with payoff
Y_t = W_t·1{W_t ≥ 200} − ∞·1{W_t ≤ 0} gives the optimal
stopping time τ* = inf{s ∈ [t,T] : W_s ≥ 200 or W_s ≤ 0} —
stop on either barrier, otherwise liquidate at T = 90.
There is no interior continuation region to exploit. [T4]
(Source: Gemini on the Snell formulation; Qwen on the
liquidate-at-day-90 conclusion)
The general optimal-stopping condition is U_t = X_t —
stop when the envelope touches the payoff process. MiniMax states it as
"when the envelope value falls below the current wealth," which is
impossible by construction: the Snell envelope is the smallest
supermartingale dominating the payoff, so
U_t ≥ X_t always. Corrected here. (See Resolved
Conflicts C-11.)
The secretary problem is invoked by two of three reports and
does not apply. Qwen and MiniMax both import it: n
rankable candidates in random order, accept-or-reject on the spot,
reject the first ⌊n/e⌋ and then take the first candidate
better than all seen, achieving
P(select the best) → 1/e ≈ 0.368; Bruss's 1/e-law (1984)
extends this to unknown n. [T2] on the
classical result — neither report supplies a resolvable citation for
Ferguson (1989) or Bruss (1984), both of which are named only through a
Wikipedia compilation. The problem concerns selecting the single
best of n rankable items under no-recall. It does
not bound the probability that a wagering sequence reaches a
wealth multiple, and MiniMax's derived "secretary bound of
P₊ ≈ 0.30" is both undomained and contradicted by MiniMax's
own fair-coin row two sections earlier. Excluded. (See Resolved
Conflicts C-12.)
The one operational item that survives. MiniMax's
practical translation is the only prescriptive content in this
subsection that the mathematics supports: pre-commit to a
stop-loss rule before the experiment begins and do not revisit it under
pressure. The justification is not that stopping improves
P₊ — under a simple random walk with absorbing barriers it
does not — but that a pre-committed rule removes the discretionary
re-entry that converts a bounded single-experiment loss into an
unbounded sequence of them. [T6] (Source: MiniMax
only)
The measurement problem, stated first. No dataset
directly measures the success rate of retail participants attempting a
2× objective at USD 100 scale over 90 days. The denominator is not
observable. What the literature supplies is a set of large-N studies of
adjacent populations that bound how unusual it is for a retail
participant to earn positive returns net of costs at all — which sets a
ceiling on P₊ for an unskilled participant, not a
measurement of it. The asymmetry is important and MiniMax states it
well: the literature supports a confident claim that the modal
retail active trader destroys capital; it does not support a confident
claim about the exact fraction achieving 2× in any given 90-day
window. [T1] on the underlying studies;
[T6] on any figure specific to this objective. (Source:
MiniMax; Gemini supplies the same studies without the caveat; Qwen
supplies neither)
Barber, B. M., & Odean, T. (2000), "Trading Is Hazardous
to Your Wealth: The Common Stock Investment Performance of Individual
Investors," Journal of Finance 55(2), 773–806. 66,465
US households at a large discount broker, 1991–1996. The most-active
traders earned 11.4% annualized against a market return
of 17.9%; the average household earned
16.4%; average annual turnover was
75%. [T1] on content. [T6] on
the DOI string: Gemini's bibliography reports 10.1111/0022-1082.00223
but never cites the paper in its body, and Gemini's DOI reliability is
demonstrably poor elsewhere (four journal/prefix mismatches plus an
internal DOI conflict between its own markdown and JSON); MiniMax
printed a speculative alternative DOI in prose, which is
excluded as a fabrication risk. The volume/issue/page citation
cross-validates across both reports and is reliable; the DOI string is
not verified in this merge. (Source: MiniMax for content, Gemini for
the DOI, Qwen invokes the paper with no figures and cites a Medium
post)
Barber, B. M., Lee, Y.-T., Liu, Y.-J., & Odean, T.
(2009), "Just How Much Do Individual Investors Lose by Trading?"
Review of Financial Studies 22(2), 609–632. Complete
trading history of all Taiwanese individual investors, 1995–1999.
Aggregate individual-investor trading losses net of costs were
approximately 2 percentage points of Taiwan market
capitalization per year. Losses concentrate among the most
active. [T1] on content; [T6] on the DOI
string (Gemini reports 10.1093/rfs/hhn046, uncited in its own body;
MiniMax leaves it unverified). (Source: MiniMax for content, Gemini
for the DOI)
Barber, B. M., Lee, Y.-T., Liu, Y.-J., & Odean, T.
(2014), "The Cross-Section of Speculator Skill: Evidence from Day
Trading," Journal of Financial Markets 18, 1–24. DOI
10.1016/j.finmar.2013.05.006 — verified. The entire Taiwanese
day-trading population, 1992–2006. Headline: less than 1% of the
day-trader population is able to predictably and reliably earn positive
abnormal returns net of fees. Cross-sectional spread:
top-decile day traders earned +37.9 bps/day after fees;
bottom-decile −28.9 bps/day. [T1] This is
the single most load-bearing empirical anchor in this cluster for an
unspecialized retail principal. (Source: MiniMax, correctly stated.
Gemini renders the same paper as "99% net unprofitable," which is a
stronger and different claim than the paper's finding about persistent,
predictable skill — pruned; see Resolved Conflicts C-13. Qwen invokes
the Taiwan study with no author, year, title, or figure.)
Chague, F., & Giovannetti, B. (2025), "The COVID-19 and
Day-Trade Pandemics in Brazil," Brazilian Review of Finance
23(1), DOI 10.12660/rbfin.v23n1.2025.94291 — verified; and Chague, F.,
De-Losso, R., & Giovannetti, B. (2019/2020), "Day Trading for a
Living?", USP Working Paper 2019_47 / FGV EESP TD 525,
RePEc:spa:wpaper:2019wpecon47. [T2] The bylines
differ by version and are reproduced as sourced: De-Losso appears on the
2019/2020 working paper but not on MiniMax's own source-list entry for
the 2025 journal article, while MiniMax's body text cites all three
authors for the 2025 paper. That discrepancy is unresolved here and the
split byline above is the conservative rendering. [T6] on
the authorship of the 2025 version. Individuals who began day trading
Brazilian equity futures 2013–2015 and persisted at least 300 days;
Gemini reports N = 19,642.
[T2][T6] on the specific percentage.[T2][T2] (Source: MiniMax)(Source: Gemini and MiniMax; Qwen invokes the Brazil study with no author, year, title, sample size, or figure)
Adjacent literature. Linnainmaa, J. T. (2011), "Why
Do (Some) Households Trade So Much?" Review of Financial
Studies 24(5), 1630–1666, DOI 10.1093/rfs/hhq108 [T1]
— a small fraction of households do most of the trading, and that group
carries prior characteristics predicting losses. Grinblatt, M., &
Keloharju, M. (2000) [T1] — comparable results in Finnish
account-level data; DOI not propagated, as MiniMax's two mentions of it
disagree with each other. Kaniel, R., Liu, S., Saar, G., & Titman,
S. (2012) [T1] — retail order imbalance predicts returns,
yet the retail investor's portfolio still loses money on net;
DOI deliberately not reproduced — MiniMax gives one
journal with a concrete DOI in a table and a different journal with an
unverified DOI in its source list, and its own digest instructs against
propagating it. The multiple-testing and replication literature —
Harvey, Liu & Zhu (2016), DOI 10.1093/rfs/hhv059; Hou, Xue &
Zhang (2020), DOI 10.1093/rfs/hhy131; McLean & Pontiff (2016), DOI
10.1111/jofi.12365 [T1] — establishes that the population
of strategies surviving multiple-testing correction and out-of-sample
replication is far smaller than the naive count, which bounds the "find
an edge" branch from above. (Source: MiniMax for the retail studies;
Gemini and MiniMax jointly for the replication trio)
Aggregation: the three reports give three different headline numbers, and two of them answer different questions.
[T6] and
explicitly labeled an extrapolation from the Taiwan <1%
anchor rather than a measurement.P(reach $200 in 90 days) = 0.18 for its best-supported
strategy (buying underpriced favorites on Kalshi/ForecastEx via zero-fee
maker limit orders), 0.10 when forced to take liquidity,
with P(ruin) > 0.98 for options buying or equity day
trading.These are not competing estimates of one quantity. Qwen and MiniMax
estimate the unconditional base rate for an unspecialized
participant with no edge; Gemini estimates the modeled
success probability of a specific strategy claiming a measured
edge. Qwen and MiniMax agree within an order of magnitude, and
their agreement — arrived at independently from the same Taiwan/Brazil
anchors — is the more defensible reading. Gemini's 0.18 is a model
output with no stated drift, volatility, trial count, code, or citation,
published in a report whose metadata simultaneously and falsely asserts
zero unverified-inference claims; it is [T6] regardless of
the [T1] label attached to it.
Merged verdict. For an unspecialized retail
participant with no measurable edge, the evidence supports a success
probability for a USD 100 → USD 200 / 90-day objective on the
order of 1% to 5%, with the strong caveat that this is an
extrapolation from adjacent populations and not a direct measurement.
[T6] on the figure; [T1] on the Taiwan and
Brazil studies that anchor it. Any estimate materially above that range
requires a demonstrated, calibrated, out-of-sample edge — and the
base-rate literature establishes that fewer than 1% of the relevant
population possesses one. [T1]
(Source: Qwen and MiniMax on the base rate; Gemini on the strategy-conditional figure, downgraded)
Established. (Recap only — each claim below
carries its tier tag and citation where it is made above; no new tags
are asserted here.) The objective is a first-passage problem, not a
return problem, and the three governing scalars are P₊,
P₀, and the law of τ₊. The relevant analytics
— two-sided barrier hitting probability, Inverse Gaussian first-passage
density, finite-horizon CDF — exist in closed form under constant drift
and diffusion. Kelly optimizes the wrong functional for this objective
and should be used as a stake ceiling, not a strategy selector. In a
subfair game, concentration maximizes P₊ and
diversification reduces it, which inverts conventional advice — and both
prescriptions are correct because they optimize different functionals,
with variance being the only mechanism supplying upper-tail mass when
E[W_T] ≤ W_0. In a game with no positive expectation, ruin
is certain in the limit. Fewer than 1% of Taiwanese day traders earn
predictable positive abnormal returns net of fees; 97% of persistent
Brazilian day traders lost money.
Not established, and no source establishes it.
(All items below are [T6] — unverified or absent from
the surveyed literature.) No validated numerical P₊,
P₀, or τ₊ distribution for this objective at
these parameters exists in any of the three reports or in any literature
they identify. The bold-versus-timid question under a hard
finite horizon is unsolved here — Dubins–Savage covers the
unbounded-time goal problem, Chen (1977) covers a discounted variant
whose directional conclusion is unverified in this merge, and the
finite-horizon HJB problem is formulated but never solved. The
Cramér–Lundberg transfer from a one-sided jump process to a two-sided
diffusion with an upper barrier is analogical and unperformed. And the
base rate for this specific objective is an extrapolation across
populations, horizons, and capital scales that no study measures
directly.
f* = (bp − q)/b; MiniMax gives
f* = p/(1+g) − q/g while claiming it simplifies to
2p − 1 at g = 1 (it yields
p/2 − q, opposite in sign at p = 0.6).
Resolution: majority rule plus arithmetic — adopted
(bp − q)/b, which is the standard result and does reduce to
2p − 1. MiniMax's version excluded; it is attached to a
correctly-verified Kelly (1956) DOI, so it carries a live risk of being
mistaken for sourced.min(x, M − x) (stake enough to reach the goal,
never overshoot); MiniMax defines it as "wagering the maximum possible
stake on every round." Resolution: majority rule — adopted
min(x, M − x). Noted that at W₀ = 100,
M = 200 the two coincide, which is why MiniMax's
operational instruction lands on the right number from the wrong
rule.[1 − (q/p)^{W₀}]/[1 − (q/p)^{M}]; MiniMax gives
1 − (q/p)^{W₀}. Resolution: Gemini's is the
classical two-barrier result. MiniMax's is the correct
M → ∞ limit — the no upper barrier case — applied
to a problem defined by an upper barrier. Adopted Gemini's.e^{2θB/σ²} for θ ≤ 0. MiniMax:
exp(−2μ_W ln2/σ_W²) when μ_W > 0, zero
otherwise. Resolution: same magnitude, opposite drift
condition. Adopted Gemini's — a diffusion with positive drift hits an
upper barrier with probability 1, so MiniMax's condition is inverted.
Gemini's 2^{(2μ/σ²) − 1} reduction verified correct.[T6] on
sourcing — Gemini attaches none, MiniMax names Karlin & Taylor
(1975) and Feller vol. 1 ch. XIV but omits both from its source
list.(1 − κ²), which implies zero growth at full Kelly.
Resolution: single-source and demonstrably wrong. Corrected to
κ(2 − κ) by inspection, tagged [T6], and
nothing built on it.ψ(u) ≤ e^{−Ru} to the ruin probability
(classical); MiniMax repurposes it as
P(reach H) ≤ e^{−R(H−x)} with no derivation.
Resolution: adopted Gemini's; MiniMax's repurposing excluded as
an unsupported inversion.R as the unique positive root and then
tabulates R ≈ −0.0002 and R ≈ −0.003.
Resolution: internally impossible — with R < 0
the bound exceeds 1 and is not a probability. MiniMax's Cramér–Lundberg
numeric table excluded entirely.τ* = inf{s : W_s ≥ 200 or W_s ≤ 0}. Resolution:
MiniMax's condition is impossible by construction (the envelope
dominates the payoff by definition). Corrected to
U_t = X_t; Gemini's degenerate two-barrier stopping time
adopted and merged with Qwen's equivalent "liquidate at day 90"
conclusion.P₊ ≈ 0.30. Resolution: the classical
1/e ≈ 0.368 result is retained as correctly stated
[T2], but the derived bound is excluded — the secretary
problem concerns selecting the best of n rankable items
under no-recall and does not bound a wealth-multiple hitting
probability, and MiniMax's own fair-coin row (P₊ = 0.50)
contradicts its 0.30 ceiling.[T6]. Gemini's N = 19,642 and 0.1% >
USD 300/day retained alongside MiniMax's top-individual USD 310/day as
compatible, distinct facts.[T6]. Gemini: 0.18
(0.10 as a taker). Resolution: these answer different
questions — the first two are unconditional base rates for an
unskilled participant, the third is a strategy-conditional model output.
Qwen and MiniMax agree within an order of magnitude and are adopted as a
1%–5% range tagged [T6]. Gemini's 0.18 is reported
separately and downgraded to [T6]: its own digest confirms
every headline probability in that report carries no drift, volatility,
trial count, code, or citation, in a report whose metadata falsely
claims zero [T6] content.P₊ is off
by roughly six orders of magnitude with the drift sign inverted). No
figure carried forward.[T6], with MiniMax's own
concession that Chen's quantitative implications for a 90-day horizon
are uncharacterized in the literature.[T1] and
the DOI string at [T6], marked unverified in this merge.
Only DOIs both reports independently support, or that one report
explicitly marks verified (Barber et al. 2014, Chen 1977, Chague et al.
2025, Gerber–Shiu 1998, Kelly 1956, Linnainmaa 2011, Harvey/Liu/Zhu
2016, Hou/Xue/Zhang 2020, McLean & Pontiff 2016), are
reproduced.[T1], replication requirement dropped; 153 of ~194 tags in
that report are [T1]). Resolution: tiers assigned
independently in this section against the spec bar rather than
inherited. Uncited mathematical formulations dropped to
[T4]/[T6]; Breiman (1961) dropped to
[T6] on citation because no source supplies a venue or DOI
and none is invented.At a USD 100 stake the binding constraint is not the minimum position size. Every vehicle in the feasible set clears its own technical floor with room to spare — USD 1.00 for fractional equities and spot crypto, USD 0.01 for Kalshi and Polymarket contracts, roughly USD 5.00 for a single low-priced listed option. (Source: Qwen, Gemini, MiniMax — unanimous.) The constraint that actually decides the experiment is the ratio of round-trip friction to stake, compounded over trade count, and behind that a small number of hard structural gates that do not bend to account size.
Two consequences follow, and they point in opposite directions from the intuition that "small accounts are cheap to run."
First, friction is scale-invariant in percentage terms for
every vehicle priced as a percentage, and scale-punitive for
every vehicle priced per contract. Kalshi's fee is linear in
contract count, so its cost as a fraction of stake is identical at USD
100 and USD 10,000 [T5] (MiniMax §4.7). But a USD
0.65 per-contract commission, a USD 0.40 round-trip fixed fee, or a
fee-rounding ceiling function all consume a fixed number of cents that
is trivial against USD 10,000 notional and material against USD 100.
Gemini's worked case is the cleanest illustration: Kalshi's ceiling
rounding turns a computed USD 0.0063 fee into a charged USD 0.01 on a
single USD 0.10 contract — a 10% drag on that stake
[T4] (Gemini, Table B).
Second, the position minimum being non-binding means
concentration is forced, not chosen. MiniMax states the
structural point directly: "at USD 100 stake, the practical MVP for a
doubling attempt is the entire stake in every candidate vehicle except
for micro-options and CME futures" [T6] (MiniMax
§2.7). A doubling attempt requires deploying substantially all
capital into a single outcome in nearly every feasible vehicle. This is
a mechanical consequence of the target multiple, not a strategy
recommendation, and it is the reason the vehicle comparison below is
conducted on friction and gates rather than on minimum size.
Denominator convention. Every total-friction figure
in Table A is expressed as a percentage of the USD 100
stake, for a full round trip (open plus close), inclusive of
deposit and on-ramp cost where one exists. This normalization is
necessary rather than cosmetic: MiniMax's source table silently mixes
three denominators — percent of stake for equities and event contracts,
percent of contract notional for standard options, percent of
net debit for spreads — which inflates the options rows by
roughly 5×–20× against the equity rows purely by denominator choice
[T6]. Gemini is the only source that carries a
deposit/withdrawal column at all; MiniMax defined deposit and withdrawal
cost into its own scope and then omitted it from its table entirely.
Totals below therefore include on-ramp cost that appears in no single
source figure.
Directional caveat on every friction number.
MiniMax's own adversarial pass concedes that its estimates "may
understate actual retail friction by 50%–200%" and "should be
interpreted as lower bounds" [T6] (MiniMax
§10.6); it also concedes no live spread tape was pulled for any
pair. Gemini's spread figures carry [T4] grey-literature
tags. No cell in Table A is a point estimate. The uncertainty is
one-sided: realized friction is more likely to exceed these figures than
to fall below them.
All percentages denominated in the USD 100 stake. Round trip = open + close. Deposit/on-ramp cost included in totals where applicable.
| Vehicle class | Min viable position (USD) | Min viable position (% of $100) | Round-trip commission | Per-contract / exchange fee | Typical spread (% notional) | Total round-trip friction (% of stake) | Feasible at $100 |
|---|---|---|---|---|---|---|---|
| Fractional equity / ETF — liquid large-cap | $1.00 | 1.0% | $0.00 at RH / Fidelity / Schwab / IBKR Lite [T5] |
$0.00 exchange; SEC/FINRA TAF ≤ $0.0001/share [T5] |
0.02%–0.25% ⁽¹⁾ | 0.02%–0.30% | Yes |
| Fractional equity — low-priced / small-cap | $1.00 | 1.0% | $0.00 [T5] |
≤ $0.0001/share [T5] |
0.50%–2.00% [T4] |
0.60%–2.10% | Yes (spread-dominated) |
| Listed options — long single leg | $5.00–$25.00 ⁽²⁾ | 5%–25% | $0.00–$0.65/contract ⁽³⁾ [T5] |
$0.06–$0.10/contract [T4] |
2.0%–10.0% [T4] |
2.5%–15.0% ⁽⁴⁾ | Marginal |
| Listed options — "micro-options" (1-share deliverable) | $0.05–$5.00 | 0.05%–5% | $0.00 claimed [T6] |
Reduced OCC fee, unquantified [T6] |
0.5%–5.0% [T6] |
Unquantified [T6] ⁽⁵⁾ |
Unverified — single source |
| Listed options — vertical debit spread | $5.00–$20.00 net debit | 5%–20% | $0.00–$1.30 per spread (2 legs) [T5] |
$0.12–$0.20 per spread [T4] |
2.0%–10.0% per leg [T4] |
5.0%–20.0% ⁽⁶⁾ | Gated (approval tier, §4.4) |
| Listed options — cash-secured put, $5 strike | $500 (strike × 100) | 500% | n/a | n/a | n/a | n/a | Infeasible ⁽⁷⁾ |
| Kalshi event contract — P ≈ 0.50 | $0.01 technical; $100 to attempt doubling | 0.01% / 100% | $0.00 | $7.00 round trip = 7.0% ⁽⁸⁾ [T4] |
1.0%–4.0% [T4] |
8.0%–11.0% | Yes technically; MA contested |
| Kalshi event contract — P ≈ 0.90 (favorite) | $0.01 technical | 0.01% | $0.00 | $1.40 round trip = 1.4% ⁽⁸⁾ [T4] |
1.0%–4.0% [T4] |
2.4%–5.4% | Yes technically; MA contested |
| Kalshi event contract — P ≈ 0.05 (longshot) | $0.01 technical | 0.01% | $0.00 | $13.30 round trip = 13.3% ⁽⁸⁾ [T4] |
2.0%–6.0% [T4] |
15.3%–19.3% | No — fee-dominant |
| ForecastEx (via IBKR Prediction Markets) | $1.00 | 1.0% | $0.00 | $0.01/contract/side; $0.02 round trip [T4] |
1.0%–3.0% [T4] |
2.0%–6.0% ⁽⁹⁾ | Yes technically; MA contested |
| IBKR CME event contracts | $0.10 | 0.1% | $0.20 round trip [T4] |
$0.20 round trip [T4] |
2.0%–5.0% [T4] |
8.0%–45.0% ⁽¹⁰⁾ | No — fixed fee dominant |
| Polymarket (USDC on Polygon) | $0.01 (1 USDC) | 0.01% | $0.00 (CLOB) [T4] |
$0.00 per contract; ≤2% on net winnings
[T6] |
0.5%–2.0% [T4] |
4.0%–10.0% ⁽¹¹⁾ | No for a MA resident (§4.5) |
| Spot crypto — advanced/pro order interface | $1.00 | 1.0% | $0.00 (fee-based, not spread-based) [T5] |
0.05%–0.60% maker/taker by volume tier [T5] |
0.10%–0.50% [T4] |
0.10%–0.60% ⁽¹²⁾ | Yes |
| Spot crypto — retail "simple trade" interface | $1.00 | 1.0% | $0.00 nominal; cost embedded in spread [T5] |
Embedded | 0.50%–2.00% [T4] |
0.80%–2.50% | Yes, ~4× the pro-interface cost |
| Crypto nano / micro futures | $20.00–$50.00 | 20%–50% | $0.20–$0.40 [T4] |
$0.10–$0.20 [T4] |
0.50%–2.00% [T4] |
3.5%–12.0% | No ⁽¹³⁾ |
| CME Bitcoin futures (standard, 5 BTC) | ~$575,000 notional; ~$200,000–$260,000 initial margin
[T6] |
~200,000% (margin) | ~$1.50–$2.50/side [T5] |
Exchange + clearing | n/a | Commission irrelevant | Decisively infeasible ⁽¹⁴⁾ |
| KalshiEX BTCPERP (perpetual) | $1.00 notional [T6] |
1.0% | Fee charged on position notional, not margin
[T5] |
Taker/maker by tier [T5] |
0.05%–0.30% [T6] |
0.25%–15.0% on margin deployed ⁽¹⁵⁾ | Marginal — leverage-binding, single source |
Notes to Table A
[T4]. Qwen asserts "<0.1%" total
friction with no decomposition and is subsumed. (Source: all
three.)[T5]. MiniMax booked the same $0.65 in
both the commission and the OCC/exchange-fee columns,
double-counting it; the per-contract regulatory and clearing charge is
separately Gemini's $0.06–$0.10 [T4], which is the figure
carried here. MiniMax's "OCC fee changed to $0.55 in 2025" is uncited
and excluded.feasible: true, contradicting its own Table A ("MARGINAL",
2.5%–12.0%); the markdown range is used.[T6] with "primary citation
withheld pending verification"; neither Qwen nor Gemini mentions the
instrument. No feasibility verdict rests on this row. MiniMax's separate
"USD 25 floor" for micro-options appears nowhere else in its own
document and is excluded.ceil(0.07 × N × P × (1−P)) per side, rounded up to the
cent. Deploying a fixed stake S at price P buys N =
S/P contracts, so the fee collapses to 0.07 × S × (1−P)
per side — on a USD 100 stake, 7 × (1−P) dollars per
side, or 14 × (1−P)% of stake round trip. The
algebra is shown so it can be checked; it appears in no source. Round
trip assumes an exit trade; whether Kalshi charges a settlement-side fee
is established in no source, so the held-to-settlement case is a ~2×
uncertainty on every Kalshi figure [T6].[T4][T6] and sits outside the friction
total.[T6]; not included in
totals.[T6]; MiniMax concedes the contract
specification was not directly retrieved. The CFTC approval itself
(release 9240-26, May 29, 2026) is separately cited. Fee is charged on
full position notional rather than posted margin, so at 5×–50× leverage
the effective drag on deployed capital is 5×–50× the notional rate.Three reports give three statements of Kalshi's fee. Qwen:
ceil(0.07 × P × (1−P) × 100)/100. Gemini:
ceil(0.07 · N · P(1−P)) [T4]. MiniMax:
0.0175 × P × (1−P) × N, with base rate 1.75%
[T6]. The two-of-three majority carries the 0.07
coefficient with a ceiling function, and the minority prunes
cleanly on quality as well as count: MiniMax states in three separate
places that it retrieved the July 7, 2026 fee schedule as an unparsed
binary and derived the 1.75% coefficient from a help-center
discussion of "expected earnings," tagging it [T6] itself.
A derived coefficient loses to two independent readings of the published
formula.
The correction is not cosmetic. It scales every Kalshi friction figure by 4×. MiniMax's headline round-trip cost at P = 0.50 is 1.75% of stake; the majority formula gives 7.0%. Gemini's independently stated Kalshi friction band of 3.00%–15.00% and Qwen's 5%–20% both bracket the corrected figure and neither brackets MiniMax's.
The more consequential finding is about shape. MiniMax characterizes the fee as "U-shaped… disproportionately punitive on contracts priced near 50%" and draws the strategy implication that the principal should "seek contracts in the tails." Both reports' own arithmetic refutes this when the denominator is the stake rather than the contract. The fee per contract is indeed maximized at P = 0.50. But a fixed stake buys N = S/P contracts, so fee-per-dollar-of-stake is 0.07 × (1−P) per side — monotonically decreasing in P. MiniMax's own nine-row table already runs monotone from 3.32% at P = 0.01 down to 0.034% at P = 0.99; the single row breaking that pattern is the one its own digest identifies as miscomputed.
The operative implication inverts: longshots are the expensive regime and favorites are the cheap one. At P = 0.05, fee alone consumes 13.3% of stake round trip; at P = 0.90, 1.4%. This matters because the longshot tail is precisely where a doubling attempt on a binary contract is mechanically available — at P > 0.50, a single contract cannot double the stake at all, since USD 100 at P = 0.90 buys 111 contracts settling at USD 111. Single-shot doubling on an event contract requires P ≤ 0.50, which is also the fee-expensive half of the curve. MiniMax's stated verdict — that event contracts are feasible "for high-confidence (>50%) outcomes… to achieve a 100% gross gain" — is arithmetically impossible and is excluded.
Cost of edge, corrected. At P = 0.50 with USD 100 deployed (200
contracts), a 5-percentage-point edge produces an expected gross gain of
USD 10. Round-trip fees of USD 7.00 consume 70% of that expected
gain — not the 17.5% MiniMax computes from its 4×-low
coefficient, and not the "needs 7%–10% edge" conclusion it draws, which
does not follow from its own numbers either way [T6].
Break-even against fee alone at P = 0.50 requires roughly a
3.5-percentage-point edge before spread.
Two gates bind hard, two bind conditionally, and one — the one most retail commentary fixates on — does not bind at all.
Pattern day trading does not bind. All three reports
agree a USD 100 account must be a cash account; the PDT regime governs
margin accounts. Gemini and Qwen recite the legacy rule as
extant; MiniMax reports it rescinded effective June 4, 2026, replaced by
intraday margin standards, citing FINRA Notice 26-10 (April 20, 2026),
SEC approval April 14, 2026 at 91 FR 20731, and quoting the Notice
directly [T5]. This is not a clean 2-of-3 majority against
MiniMax: Qwen's citation is fabricated ("SEC Rule 2222" for a FINRA
rule) and its description of the rule is inverted, and Gemini reciting
the legacy citation is not a dated claim about currency. MiniMax is the
only source making a dated claim, and it is the better-cited one — a
Federal Register number, a notice number, and a block quotation. It
remains unverified against primary source, and two
siblings contradict it. The row closes regardless: a USD 100
cash account is outside the PDT regime under either version of the
rule.
T+1 settlement binds, and it is the real trade-count
ceiling. Gemini and MiniMax both cite SEC Rule 15c6-1 / 17 CFR
§ 240.15c6-1 [T5]; MiniMax dates the amendment to May 28,
2024. Qwen asserts T+2 three times and is pruned as stale for a
2026-dated report — a material error, since T+2 roughly halves the
achievable round-trip frequency the report then models. Under T+1,
unsettled sale proceeds cannot fund the next purchase, yielding roughly
one round trip per two business days on a single security, or
~30 round trips across a 90-day window under best-case timing
[T5] (MiniMax).
Free-riding binds, with an unresolved trigger count.
Selling a security purchased with unsettled funds is a Regulation T
cash-account violation (12 CFR Part 220) carrying a 90-day cash-up-front
restriction. Gemini states three good-faith violations trigger the lock;
MiniMax states three in a 12-month period in one table row and one
violation in an adjacent row of the same table. Both reports' authority
citations are wrong — Gemini attributes it to "FINRA Rule 2210"
(Communications with the Public) and MiniMax lists "FINRA" as the
authority for a Federal Reserve Board regulation. The regulation is
cited at part level here; the 1-versus-3 trigger count is
[T6] and is decision-relevant because it directly
caps maximum trade count.
Options approval binds conditionally, and this is the
sharpest live disagreement. All three agree long single-leg
calls and puts are available to a funded USD 100 account at the lowest
or second-lowest tier [T5]. On multi-leg spreads they
split: Gemini and Qwen say spreads require Tier 3/4, which mandates a
margin account with USD 2,000 minimum equity, so a USD 100 account is
capped at single legs; MiniMax says a debit spread with maximum loss ≤
USD 100 is structurally possible in a cash account. These are not
actually contradictory — they answer different questions. A defined-risk
debit spread is cash-securable in principle; the obstacle is
that the approval tier permitting it is one a USD 100 account is
unlikely to clear, and MiniMax itself notes brokers commonly impose
30 days of account seasoning for Level 3 and 60 days for Level
4 [T6]. Against a 90-day clock, a 30-day seasoning
requirement consumes a third of the window. The unified finding:
structurally cash-securable, practically gated.
KYC and funding latency bind on the clock, not the
capital. No venue in the feasible set requires more than USD 1
to open — Robinhood, Schwab, Fidelity, and IBKR Lite all set USD 0,
Kalshi and Coinbase approximately USD 1 [T6]
(MiniMax). The cost is time. Gemini quantifies ACH funding
holds at 3–5 business days, "5.5% of the 90-day clock," citing 31 CFR §
1020.220 [T5]; MiniMax gives per-venue onboarding of
instant-to-1 business day at Robinhood, Kalshi, and Polymarket, 1–3 days
at Schwab and Fidelity, and 1–5 days at IBKR and Coinbase with manual
review running 1–4 weeks, citing FINRA Rule 2090 and 31
CFR § 1010.230 [T5]. The two are consistent and additive:
identity verification and funds availability are separate waits. A
conservative reserve of 5 business days before the account is
tradeable is the merged figure, with a low-probability tail to
four weeks if a manual review is triggered. The 90-day clock starts at
tradeability, not at application.
| Gate | Authority | Rule citation | Applies to a $100 account? | Practical consequence | Primary source URL |
|---|---|---|---|---|---|
| Pattern-day-trader minimum equity | FINRA / SEC | FINRA Rule 4210(f)(8)(B) (legacy) [T5]; reported
rescinded eff. June 4, 2026 — FINRA Notice 26-10 (Apr 20, 2026), SEC
approval Apr 14, 2026, 91 FR 20731 [T5], unverified |
No — cash account is outside the regime under either version | Legacy $25,000 floor and the 4-day-trades-in-5 designation never
applied to a cash account. Not the binding constraint. Broker phase-in
reported through Oct 20, 2027 [T6] |
https://www.finra.org/rules-guidance/notices/26-10 |
| T+1 cash settlement | SEC | SEC Rule 15c6-1 / 17 CFR § 240.15c6-1, amendment eff. May 28, 2024
[T5] |
Yes — all cash accounts regardless of size | Unsettled proceeds cannot fund the next purchase. ~1 round trip per 2 business days; ~30 round trips per 90 days best case. Daily turnover capped at the stake | https://www.sec.gov/rules/final/34-96930.pdf |
| Free-riding / good-faith violation | Federal Reserve Board (not FINRA) | Regulation T, 12 CFR Part 220 [T5]; pinpoint disputed
across sources |
Yes | 90-day cash-up-front account restriction. Trigger count
unresolved: 1 violation (MiniMax, one row) vs 3 in 12 months (Gemini,
MiniMax other row) [T6]. Directly caps maximum
trade count |
https://www.sec.gov/investor/alerts/cashaccounts.pdf |
| Options approval — long single leg (Level/Tier 1–2) | FINRA / OCC | FINRA Rule 2360 [T5] |
Yes, and clearable | Granted at or near funding by IBKR, Schwab, Robinhood. Long calls and puts accessible at $100 | https://www.finra.org/rules-guidance/rulebooks/finra-rules/2360 |
| Options approval — multi-leg spreads (Level/Tier 3–4) | FINRA / OCC / broker | FINRA Rule 2360 [T5]; broker tier policy
[T6] |
Yes — and likely blocking | Gemini and Qwen: Tier 3/4 requires a margin account with $2,000 minimum equity. MiniMax: a debit spread with max loss ≤ $100 is cash-securable. Unified: structurally permissible, practically gated by an approval tier a $100 account is unlikely to clear | https://www.finra.org/rules-guidance/rulebooks/finra-rules/2360 |
| Account seasoning for options tiers | Broker policy under FINRA Rule 2360 | Supplementary material, pinpoint unverified [T6] |
Yes if Level 3+ is sought | ~30 days for Level 3, ~60 days for Level 4 at some
brokers [T6]. Against a 90-day window this consumes 33%–67%
of the clock before the strategy is available |
https://www.finra.org/rules-guidance/rulebooks/finra-rules/2360 |
| Kalshi taker fee formula (non-linear + ceiling) | Kalshi (DCM exchange rule) | Fee schedule, July 7, 2026 [T4] |
Yes — every Kalshi trade | ceil(0.07 × N × P × (1−P)) per side. On a fixed stake
this is 7 × (1−P) dollars per side per $100,
monotonically decreasing in P. 7.0% of stake round trip at P =
0.50; 13.3% at P = 0.05; 1.4% at P = 0.90. Ceiling rounding
alone costs 10% on a single $0.10 contract |
https://kalshi.com/docs/kalshi-fee-schedule.pdf |
| Polymarket / Kalshi federal DCM authorization | CFTC | CFTC order Jan 3, 2022 (Polymarket, $1.4M, failure to register as
SEF) [T5]; Amended Order of Designation reported Nov 2025
[T6] — Wikipedia-sourced only |
Disputed — see §4.6 | Federal gate reported open following the QCEX acquisition (MiniMax, Qwen). Gemini reports the 2022 consent order and US IP geo-blocking still governing. No CFTC release number is cited by any source for the Amended Order | Only source cited by any report: https://en.wikipedia.org/wiki/Polymarket |
| Massachusetts state gate — non-sports event contracts | MA Superior Court; MA SJC (pending); MA AG | Commonwealth v. KalshiEX LLC, preliminary injunction Jan
2026 [T6] (no docket number given by any source); M.G.L. c.
23K [T5] |
Yes — and this is the operative gate | All three reports converge: a MA resident should not assume lawful access to non-sports event contracts. MA SJC has not ruled; CFTC has filed an amicus asserting federal preemption; no venue has published an explicit MA policy for non-sports contracts | Only source cited: https://en.wikipedia.org/wiki/Kalshi |
| Massachusetts geofence — sports contracts | MA Superior Court order | Commonwealth v. KalshiEX LLC PI, Jan 2026
[T6] |
Yes | KalshiEX required to geofence MA residents from sports markets. Out of scope for the economic/macro contract universe but establishes the state's legal theory | https://en.wikipedia.org/wiki/Kalshi |
| Minimum account funding | Each venue | Venue terms [T6] |
No — not binding | Robinhood, Schwab, Fidelity, IBKR Lite, IBKR Prediction Markets: $0. Kalshi, Polymarket, Coinbase: ~$1. CME futures via IBKR: $2,000 margin account. A $100 stake clears every venue except futures | Each venue's account-opening page |
| KYC / customer identification | FinCEN; FINRA | 31 CFR § 1020.220 [T5]; 31 CFR § 1010.230
[T5]; FINRA Rule 2090 [T5] |
Yes — every venue | Requires SSN/ITIN, address, employment, financial profile. Online onboarding 1–5 business days; manual review 1–4 weeks. Bottleneck is latency, not capital | https://www.finra.org/rules-guidance/rulebooks/finra-rules/2090 |
| ACH funding latency | FinCEN / broker | 31 CFR § 1020.220 [T5] |
Yes | Initial ACH holds of 3–5 business days = ~5.5% of the 90-day clock. Merged reserve: 5 business days before the account is tradeable. The 90-day clock starts at tradeability, not application | https://www.ecfr.gov/current/title-31/section-1020.220 |
This was flagged as the highest-stakes open question in the commissioning prompt. It resolves into three separable propositions with very different evidentiary standing, and reporting them as one status label would misrepresent all three.
(a) Lawful Massachusetts access to non-sports event contracts
— all three reports converge on "no, or contested." This is the
operative finding. Qwen: "a Massachusetts resident cannot
currently participate in these federally regulated markets without
facing potential legal jeopardy," citing preliminary injunctions against
Kalshi, a Suffolk County Superior Court order requiring geofencing, and
30+ active prediction-market lawsuits nationwide [T4].
MiniMax: MA access to non-sports event contracts is "contested"; the
Commonwealth v. KalshiEX injunction reaches Polymarket only by
analogous state-law argument; the MA SJC has not ruled; no venue has
published an explicit MA policy for non-sports contracts; "the
conservative reading is that MA residents should not assume they have
lawful access… until the MA SJC rules, the CFTC prevails in the
litigation, or each venue publishes an explicit MA policy"
[T6]. Gemini: "non-compliant for US/MA retail participants"
[T6]. Three independent reports, three different routes,
one conclusion. For a Massachusetts-resident principal,
Polymarket is not an available vehicle, and the same reasoning reaches
Kalshi and ForecastEx on non-sports contracts. This is the
finding that carries consequence, and it does not depend on resolving
(b) or (c).
(b) Current US retail accessibility — genuinely contested,
and the disagreement is not reconcilable from these sources.
MiniMax states Polymarket unblocked US customers on December 2,
2025, operating on USDC on Polygon. Gemini states Polymarket is
geo-blocked to US IPs under the 2022 CFTC consent order and grades it
NOT FEASIBLE on that basis. Qwen sits between them: US operations were
"halted by CFTC enforcement, later resumed via acquisition of a licensed
entity" [T4]. On a head count, Qwen and MiniMax agree that
operations resumed and Gemini is the minority; on citation quality,
Gemini's basis is a 2022 order that predates the events the other two
describe, which makes it stale rather than contradictory. The
direction — resumption via acquisition of a CFTC-licensed entity —
carries two-of-three support and the better temporal footing. The
specific date of December 2, 2025 is single-source and
Wikipedia-sourced, and is [T6].
(c) Federal DCM authorization — reported, but the citation
does not support the tier it was given. MiniMax reports
Polymarket acquired QCEX, a CFTC-licensed derivatives exchange and
clearinghouse, for USD 112 million in 2025; that DOJ
and CFTC ended their investigations on July 15, 2025
without new charges; and that the CFTC issued an Amended Order
of Designation in November 2025. Qwen independently
corroborates the acquisition-of-a-licensed-entity mechanism without
dates. MiniMax tags all six chronology items [T5] — its own
tier for primary regulatory documentation — while sourcing every one of
them solely to https://en.wikipedia.org/wiki/Polymarket, a
tertiary source. The contrast that makes this a downgrade rather than a
quibble: the same document cites CFTC press releases by number
three times elsewhere (9240-26, 9249-26, 9267-26) and cites none for the
Amended Order, nor a Federal Register entry, nor the CFTC's DCM
registry. The chronology is retained; the tier is downgraded to
[T6]; and the "federally authorized — settled" verdict is
not carried forward without primary CFTC confirmation.
The asymmetry worth stating plainly: the federal question (c) is the one with the weakest sourcing and the least consequence for a Massachusetts principal, while the state question (a) is the one with three-way convergence and all of the consequence. Resolving the federal gate open would not open the state gate.
Ranking by total round-trip friction as a share of stake, which is the only axis on which the three reports produce comparable, checkable numbers:
[T5]. Unanimous across all three
reports. Its constraint is not cost but the absence of any structural
mechanism to double: 1:1 exposure, no leverage without borrowing, and
T+1 capping turnover at ~30 round trips over the window.Aggregated over trade count, friction becomes the dominant term. MiniMax projects Kalshi at 50 round trips as consuming more than the entire stake; at the corrected 0.07 coefficient the projection is worse still, though the linear model overstates it because drag is capped at the stake and the stake decays. The directional result survives the arithmetic objection: at any meaningful trade count in a per-contract-fee vehicle, friction alone is ruinous absent a persistent positive edge. Fractional equity at 0.02%–0.30% per round trip is the only vehicle where 50 round trips costs less than 15% of stake.
None of this addresses whether any strategy in any of these vehicles carries positive expected value, which is the subject of later sections. Feasibility is a necessary condition, not a sufficient one; a vehicle can be perfectly accessible and still have no available edge.
ceil(0.07 × N × P(1−P)); MiniMax gives
1.75%. Resolved to 0.07 on majority and on
quality — MiniMax self-tags its coefficient [T6] and states
it was derived from a help-center article, not read from the fee
schedule. Scales all Kalshi friction figures by 4×.[T6] — a ~2×
uncertainty on every Kalshi figure.[T6].[T5] while sourcing
it solely to Wikipedia, in a document that cites three CFTC release
numbers elsewhere. Downgraded to [T6];
chronology retained, "settled" verdict withheld pending primary CFTC
confirmation.[T6]. Both
reports' authority citations are wrong (Gemini: "FINRA Rule 2210";
MiniMax: "FINRA / Reg T § 220.4" for a Federal Reserve regulation);
cited at part level as 12 CFR Part 220.feasible: true, 7.5%) contradicts its markdown
("MARGINAL", 2.5%–12%); markdown used.[T6] with "primary citation withheld pending verification";
no sibling corroboration. Included as an explicitly
single-source unverified row carrying no feasibility verdict.
MiniMax's separate "$25 floor" figure (appears nowhere else in its own
document) and its "$480 max loss on a debit spread" (wrong — max loss on
a long vertical is the net debit) both excluded.[T1] to include primary regulatory statutes, tagging FINRA
Rule 4210, 17 CFR § 240.15c6-1, and 31 CFR § 1020.220 as
[T1]. Every Gemini regulatory citation re-tiered to
[T5] under the CASINO taxonomy before import; tag
counts are not aggregated across reports.This section grades every strategy for which any of the three source reports claimed peer-reviewed support. It is organized around a single question: does the published literature contain a documented effect large enough, fast enough, and cheap enough to move USD 100 to USD 200 inside 90 days? The answer, developed below and summarized in Table C, is no — and the more important finding is why the answer is no, because the failure mode is structural rather than a matter of picking the wrong anomaly.
A methodological warning governs the whole section. This cluster
carried the highest citation-fabrication rate of any part of the merged
corpus. Independent verification of MiniMax's Cluster 3 bibliography
found roughly one third of its distinct citations fabricated or
materially misattributed, including its foundational
volatility-risk-premium citation, the sole [T1] citation
supporting its entire merger-arbitrage section, and a paper attributed
to "Penn (2025)" that does not exist (Source: MiniMax digest
§5.1). Gemini's citation apparatus failed differently but no less
badly: its own self-declared "Chain-of-Verification Audit" is disproven
from inside its own deliverable, and every headline probability it
reports is uncited and unmodeled (Source: Gemini digest §7.0,
§7.2). Qwen's Cluster 3 attributes post-earnings drift to a paper
about IPO underperformance and short-term reversal to a paper about
momentum (Source: Qwen digest §7c). Accordingly, several
figures that appear in one source report with confident precision are
downgraded or struck here, and the reasoning for each is recorded in
Resolved Conflicts at the end. Where a claim survives
on one source only and could not be corroborated, it is tagged
[T6] and labeled as such rather than laundered into
apparent consensus.
Before evaluating any specific strategy, the base rate on published anomalies has to be set, because it determines how much weight a single published effect size deserves.
Multiple testing. Harvey, Liu & Zhu (2016),
"…and the Cross-Section of Expected Returns," Review of Financial
Studies 29(1), 5–68, DOI 10.1093/rfs/hhv059, evaluate
roughly 315 published factors and show that the conventional
t > 2.0 significance bar, applied across a literature
that has run hundreds of tests, produces a false-discovery rate well
above tolerance. Their corrected threshold is t > 3.0
[T1]. (Source: Gemini, MiniMax; Qwen credits Harvey and
co-authors with the correction but never states the threshold or the
factor count.) The practical consequence is blunt: a published
anomaly reporting t ≈ 2.2 carries roughly the evidentiary
weight of an unpublished one.
Replication. Hou, Xue & Zhang (2020),
"Replicating Anomalies," Review of Financial Studies 33(5),
2019–2133, DOI 10.1093/rfs/hhy131, re-implement 452
anomalies under a common q-factor framework. 65%+ fail to
replicate at t ≥ 1.96; 82%+ fail at
t ≥ 2.78 [T1]. The mechanism they
identify matters more than the headline: the failures concentrate in
equal-weighted microcap portfolios, i.e. precisely the corner of the
universe where an anomaly's paper alpha is destroyed by transaction
costs (Source: Gemini, MiniMax). The same authors' q-factor
model — Hou, Xue & Zhang (2015), "Digesting Anomalies," RFS
28(3), 650–705, DOI 10.1093/rfs/hhu068 [T1] —
absorbs most of the apparent alpha in the anomaly literature into four
factors, meaning many "anomalies" were factor exposure wearing a costume
(Source: MiniMax).
Post-publication decay. McLean & Pontiff (2016),
"Does Academic Research Destroy Stock Return Predictability?"
Journal of Finance 71(1), 5–32, DOI
10.1111/jofi.12365, study 97 anomalies and find returns
decay 58% post-publication, decomposed as 26%
academic over-fitting plus 32% arbitrage exploitation
[T1] (Source: Gemini, MiniMax). MiniMax reports
the same 58% headline with the underlying levels — pre-publication
in-sample alpha ≈0.84%/month falling to ≈0.34%/month out-of-sample
post-publication (Source: MiniMax). This is the single most
important calibration number in the section: any effect size
quoted from a paper should be haircut by roughly 58% before it is used
as a forward expectation, and by more than that if the strategy
was crowded after publication.
Publication survivorship. Linnainmaa & Roberts
(2018), "The History of the Cross-Section of Stock Returns,"
RFS 31(7), 2606–2649, DOI 10.1093/rfs/hhy002
[T1], show that a substantial fraction of apparent
discovery is survivor bias in what gets published, so the anomaly
literature systematically overstates effect sizes even before decay
(Source: MiniMax).
The dissent, which none of the three reports honestly
surfaced. Jensen, Kelly & Pedersen (2023), Journal of
Finance 78(5), reach the opposite conclusion from Hou-Xue-Zhang:
under a hierarchical Bayesian treatment most factors do
replicate, and there is no replication crisis in finance
[T1]. MiniMax cites this paper approvingly elsewhere in its
cluster while suppressing its headline finding — the finding that most
directly contradicts MiniMax's own thesis (Source: MiniMax digest
§5.3, flag M2). Gemini and Qwen do not mention it at all. This
section carries the dissent explicitly. The honest reading is that
the replication literature is contested, not settled,
and the asymmetry runs as follows: if Hou-Xue-Zhang are right, published
effect sizes are mostly noise; if Jensen-Kelly-Pedersen are right, the
effects are real but small, priced, and crowded. Neither branch
produces a 100% return in 90 days at USD 100. The dispute is
therefore load-bearing for academic finance and irrelevant to this
objective — a useful result, because it means the conclusion below is
robust to the outcome of the biggest open argument in the field.
MiniMax states that "roughly 80–90% of published anomalies either
fail to replicate out-of-sample, decay materially post-publication, or
are subsumed by canonical factor models" [T6]. That figure
is uncited, exceeds Hou-Xue-Zhang's verified 65% at the conventional
threshold, and is not adopted here. The defensible statement is
65% failure at t ≥ 1.96, 82% at
t ≥ 2.78, plus 58% decay among those that do
replicate [T1].
Applying the above, the factors with genuine multiple-testing-adjusted, out-of-sample support are few. All citations below are the corrected forms; several appear in the source reports with wrong journals or wrong DOIs and are repaired in Resolved Conflicts.
Robust tier [T1]: market excess return
— Sharpe (1964), JF 19(3), 425–442, DOI
10.1111/j.1540-6261.1964.tb02865.x; Fama & French
(1993), JFE 33(1), 3–56, DOI
10.1016/0304-405X(93)90023-5. Profitability (RMW) —
Novy-Marx (2013), "The Other Side of Value: The Gross Profitability
Premium," JFE 108(1), 1–28, DOI
10.1016/j.jfineco.2013.01.003. Investment (CMA) — Cooper,
Gulen & Schill (2008), JF 63(4), 1619–1663, DOI
10.1111/j.1540-6261.2008.01369.x; Titman, Wei & Xie
(2004), JFQA 39(4), DOI 10.1017/S0022109000003125.
Short-term reversal — Jegadeesh (1990), JF 45(3), 881–898, DOI
10.1111/j.1540-6261.1990.tb05110.x; Lehmann (1990), "Fads,
Martingales, and Market Efficiency," JFQA 25(1), 1–21, DOI
10.2307/2330889. Momentum — Jegadeesh & Titman (1993),
JF 48(1), 65–91, DOI
10.1111/j.1540-6261.1993.tb04702.x; Asness, Moskowitz &
Pedersen (2013), "Value and Momentum Everywhere," JF 68(3),
929–985, DOI 10.1111/jofi.12021. Carry — Koijen, Moskowitz,
Pedersen & Vrugt (2018), JFE 127(2), 197–225, DOI
10.1016/j.jfineco.2017.11.002. Long-term reversal — DeBondt
& Thaler (1985), JF 40(3), 793–805, DOI
10.1111/j.1540-6261.1985.tb05004.x. (Source: MiniMax,
with DOIs verified in its own audit pass.)
Contested tier [T2]: value (HML) —
premium materially reduced over 2017–2020 with a partial rebound in
higher-inflation regimes post-2021, and rejected as an independent
factor by the q-model once investment and profitability are included
(Source: MiniMax); Fama & French (2015), JFE
116(1), 1–22, DOI 10.1016/j.jfineco.2014.10.010. Quality
(QMJ) — Asness, Frazzini & Pedersen (2019), "Quality Minus Junk,"
Review of Accounting Studies 24(1), DOI
10.1007/s11142-018-9470-2 [T1] for the effect,
[T2] for its independence. Betting Against Beta — Frazzini
& Pedersen (2014), JFE 111(1), 1–25, DOI
10.1016/j.jfineco.2013.10.005 [T1], with the
caveat that at least one subsequent paper argues BAB is subsumed by
standard risk factors; that subsuming citation appears in one source
only and could not be verified, so the subsumption claim is carried at
[T6].
Decayed tier [T2]: accruals — Sloan
(1996), Accounting Review 71(3), 289–315 — weakened materially
post-publication and partially subsumed by profitability. Net stock
issuance — Loughran & Ritter (1995) — material decay since
publication (Source: MiniMax).
Note what the robust tier contains: long-horizon,
cross-sectional, many-name premia measured in tenths of a percent per
month. Profitability runs ~0.5%/month, investment ~0.3%/month,
quality ~0.4%/month, low-volatility ~0.4%/month at the levels MiniMax's
Table C reports [T6] — those specific monthly figures are
single-sourced and internally inconsistent with MiniMax's own prose in
places, so they should be read as order-of-magnitude, not as
measurements. The order of magnitude is the point. Author
derivation [T6]: a 0.4%/month premium
compounds to roughly 1.2% over 90 days, against an objective requiring
100% — a shortfall of a factor of roughly eighty, not a factor
of two or three.
MiniMax imposes five structural gates, and they are the most useful
analytical contribution in its cluster because they convert a return
question into a feasibility question (Source: MiniMax §11.1)
[T6] — the gates are the report's own construction, not a
literature finding, and are labeled as such:
t ≈ 2.0 is attainable; with 15, only very
large effects are detectable.Gate 2 is where the factor literature dies at this scale. Every
factor premium above is a cross-sectional claim: it is
the average return spread across a wide portfolio, and its documented
Sharpe ratio assumes 50–200 names rebalanced monthly at institutional
size. MiniMax's formulation is exact and worth preserving verbatim:
"Single-name concentration at USD 100 bet is equivalent to a binary
bet, not a factor strategy" [T6]. At USD 100 spread
across ten names, each position is USD 10; at 30 names, USD 3.33.
Neither is a portfolio — the first is under-diversified relative to the
effect being harvested, and the second is dominated by per-trade market
impact.
Gate 1 fails independently. MiniMax computes effective retail
friction on a ten-name momentum rotation at 4–8% per round
trip once USD 0.10–0.50 of market impact per trade is charged
against a USD 10 position, even at zero commission (Source: MiniMax
§11.2) [T6]. That figure is eight to sixteen times its
own stated 0.5% budget. Its Table C nonetheless grades momentum
"Marginal"; on its own arithmetic the grade must be "No," and Table C
below reflects the corrected grade.
Gemini reaches the same structural conclusion from a different direction: cross-sectional momentum is graded NOT feasible because "90 days is too short for factor realization," short-term reversal NOT feasible because turnover spread drag consumes the alpha, index reconstitution NOT feasible on calendar grounds, VRP NOT feasible on margin grounds, merger arbitrage NOT feasible on minimum-capital grounds (Source: Gemini Table C). Qwen concurs qualitatively: the factor literature is "barren at this capital scale" (Source: Qwen).
All three reports independently conclude that no documented factor-based anomaly is cleanly executable at USD 100 within 90 days after friction. This is the strongest three-way agreement in the entire cluster, and it is the finding on which the merged report should place the most weight — not because the sources are individually reliable (they are not), but because they reach it through three non-overlapping chains of reasoning: MiniMax through friction arithmetic, Gemini through per-strategy capital minimums, Qwen through diversification requirements.
The short-horizon literature is where a 90-day window could in principle bind, so each documented effect is worked through effect size, decay, minimum capital, and friction breakeven.
Evidence. Ball & Brown (1968), Journal of
Accounting Research 6(2), 159–177, DOI 10.2307/2490232
[T1]; Bernard & Thomas (1989), JAR 27
Supplement, 1–36, DOI 10.2307/2491256 [T1];
Bernard & Thomas (1990), Journal of Accounting and
Economics 13(4), 305–340, DOI
10.1016/0165-4101(90)90008-R [T1]. Prices
drift in the direction of the earnings surprise, sorted on standardized
unexpected earnings (SUE), for up to 60 trading days after the
announcement (Source: MiniMax, Gemini; Qwen describes the phenomenon
correctly but attributes it to Loughran & Ritter (1995), a paper
about IPO long-run underperformance — misattribution,
discarded).
Effect size. The two quantified sources disagree on
horizon rather than on substance. MiniMax reports a top-vs-bottom SUE
decile spread of ≈7–10% over the 60-day window in original samples,
reduced to 3–5% in recent samples [T2]. Gemini reports
+2.0% to +5.0% over a 30-day hold on top-decile SUE,
with ~35% post-2000 decay [T2]. Adjusting MiniMax's 60-day
figure to a 30-day horizon brings the two into rough agreement;
MiniMax's own verification pass separately flags the 7–10% original
figure as likely inflated against Bernard & Thomas's canonical hedge
return. Adopted: +2% to +5% on the long leg over a 30-day hold,
post-decay [T2]. Both reports' specific ranges are
report harmonizations rather than figures lifted from the papers, which
is why this is graded [T2] and not [T1].
Decay. 35% post-2000 (Gemini) to ~50% (MiniMax)
[T2]. MiniMax's decay citation does not support the claim —
the paper it cites studies profitability and book-to-market, not PEAD
(Source: MiniMax digest flag F20/M4) — so the ~50% figure loses
its support and the range narrows toward Gemini's 35%, itself uncited.
Treat the decay as real, directionally large, and imprecisely
measured.
There is a substantive mechanism claim worth carrying: a follow-up
working paper suggests modern PEAD is now concentrated in stocks with no
sell-side analyst following, making the residual a
limited-attention premium rather than an
earnings-processing premium [T3], single-sourced
(Source: MiniMax). If true, it implies the residual anomaly
lives in exactly the low-liquidity names where retail friction is
worst.
Feasibility. Gemini grades PEAD
feasible at $5.00 minimum on commission-free
fractionals, executing across 10–15 earnings events in 90 days with SUE
> 2.0 standard deviations, 30-day holds, max 5 positions (Source:
Gemini Table C and JSON strategies[1]). MiniMax grades
it marginal, requiring friction under 0.5% per round
trip on sub-5bp-spread S&P 500 names, and computes the realistic
single-name outcome as a 5–15% gain — USD 105 to USD 115 on a USD 100
stake. Both are consistent with the same conclusion: PEAD is the
most executable documented effect in the corpus and is still an order of
magnitude short of the target. Author
derivation [T6]: at +2% to +5% per 30-day event
with a maximum of three sequential holds in 90 days, the compounded
range is +6% to +16% before friction. Reaching +100% would require
either leverage the account cannot obtain or a concentration bet whose
outcome is driven by the single name's idiosyncratic variance, not by
PEAD.
Evidence. Jegadeesh (1990) and Lehmann (1990), cited
above [T1]. Past one-week losers outperform over the
subsequent one-to-five days; the economic interpretation is compensation
for supplying liquidity (Source: MiniMax, Gemini).
Effect size. MiniMax: ~0.4–0.6% per week, and
separately 1–2% over 1–5 days. Gemini: +0.5% to +1.0%
weekly [T2]. Two of three sources converge on
~0.5%/week gross; adopted. Qwen offers no figure and misattributes the
effect to Jegadeesh & Titman (1993), which documents
intermediate-horizon momentum — the opposite phenomenon (Source:
Qwen digest §7c); discarded.
Decay and mechanism. The effect concentrated
historically in small illiquid stocks and has compressed as electronic
market-making absorbed the liquidity-provision return [T2].
MiniMax attributes the decay to HFT market-making but cites a 2001 paper
for it, which cannot document a cause that post-dates it (Source:
MiniMax digest flag M3); the attribution is retained as a mechanism
hypothesis at [T6], the decay itself at
[T2].
Friction. This is the decisive column. Gemini
charges 80%+ of the gross effect to fees and sets the
friction breakeven at $500 — five times the available
stake — grading it NOT feasible (Source: Gemini
Table C). MiniMax's prose computes a total friction budget of
4–12% over 90 days across ~12 round trips, "eating most
of the documented effect size," while its own Table C simultaneously
reports "~6% net in 90d" and grades it "Marginal" — an unreconciled
internal contradiction (Source: MiniMax digest flag N5). Taking
MiniMax's prose over its table (the prose shows its work), and combining
with Gemini — author derivation [T6]:
gross of 0.5–1.0%/week over 13 weeks is 6.5–13%, against friction of
4–12% over the same window, bounding expected net at −5.5%
(worst gross against worst friction) to +9% (best gross against best
friction), with the full cross-range running −11.5% to +9%.
The expected net straddles zero on any pairing. Graded
No.
Evidence. Jegadeesh & Titman (1993) and Asness,
Moskowitz & Pedersen (2013) for cross-sectional momentum
[T1]; Moskowitz, Ooi & Pedersen (2012), "Time Series
Momentum," JFE 104(2), 228–250, DOI
10.1016/j.jfineco.2011.11.003 [T1] for the
long-only time-series variant (Source: MiniMax). Gemini
attributes cross-sectional momentum to Harvey, Liu & Zhu (2016) — a
multiple-testing critique, not a momentum result, a misattribution
Gemini's own audit identifies (Source: Gemini digest §7.1 item
8); discarded and replaced with the correct citations.
Effect size. MiniMax: zero-cost long-short earns
~1%/month over 1965–1989, falling to ~0.5%/month post-publication, gross
Sharpe ≈0.4 falling to net ≈0.1–0.2 after realistic costs. Gemini:
+4% to +8% annualized excess return with 58% post-publication
decay. These are compatible — 0.5%/month is ~6%/year, inside
Gemini's range. Adopted: 4–8% annualized post-decay
[T2]. Momentum also carries documented crash risk: MiniMax
flags the 2008–2009 disaster period, where the strategy's return
distribution is severely left-skewed [T1].
Robustness dissent. One source cites a paper arguing
the 12-month momentum premium is not robust across asset classes and
time periods under multiple-dataset bootstrap inference; that citation
could not be verified in a single pass, though its venue is poorly
indexed and a miss is weak evidence. Carried at [T6],
single-sourced (Source: MiniMax digest flag F21).
Feasibility. Three of three: No. Gemini — 90 days is too short for factor realization at monthly rebalance. MiniMax — effective retail friction 4–8% per round trip against a 0.5–1% budget. Qwen — momentum requires holding a basket across monthly or quarterly horizons, incompatible with both the deadline and rapid deployment. A 90-day, ten-name, long-only momentum sleeve returns an expected 2–5% with substantial tail risk on MiniMax's own numbers.
Evidence. Carr & Wu (2009), "Variance Risk
Premiums," RFS 22(3), 1311–1341, DOI
10.1093/rfs/hhn038 [T1] (Source: MiniMax,
DOI verified in its own audit; Gemini cites the same paper with a
Journal of Banking & Finance DOI prefix, which its own audit flags
as mismatched — MiniMax's DOI adopted). Coval & Shumway (2001),
"Expected Option Returns," JF 56(3), 983–1009, DOI
10.1111/0022-1082.00352 [T1] (Source:
MiniMax, verified). The core claim — implied volatility
systematically exceeds subsequently realized volatility, so option
sellers earn a premium — is asserted by all three reports
[T1].
Effect size. This is the worst-corroborated number
in the section. MiniMax reports the VRP at ~0.10% per
day in its prose and ~0.5% per day gross in
its Table C — a five-fold internal contradiction inside one document
(Source: MiniMax digest flag N2). It also attributes a
short-straddle Sharpe of 0.50–0.75 to Coval & Shumway (2001);
verification found that paper's headline result is that zero-beta
straddles earn large negative returns of roughly −3%/week, and
the 0.50–0.75 figure appears invented (Source: MiniMax digest flag
M5). Both MiniMax magnitudes and the Coval-Shumway Sharpe
are struck. Gemini reports +1.0% to +2.0% monthly,
structurally stable, which is a report harmonization rather
than a paper figure but is at least internally consistent and not
contradicted. Adopted with a wide error bar: +1% to +2% per
month gross to the short-volatility seller [T2],
single-sourced.
The economically correct framing, which two of three sources state,
is that the VRP is compensation for bearing crash risk, not a
free lunch [T1]. MiniMax records the canonical
realization: short-volatility retail positions lost 30–90% of portfolio
value in a single day on 5 February 2018, and comparable losses occurred
in March 2020 [T4]. A strategy with a positive mean and a
left tail that can remove 90% of capital in one session is, under a
first-passage objective, worse than its Sharpe ratio suggests — ruin is
absorbing, and the strategy's whole edge accrues in the 95% of paths
that do not matter if the 5% path ends the experiment.
Minimum capital. Two of three agree the binding
constraint is $2,000 — the margin floor for
cash-secured puts, and the practical floor for short-option approval
(Source: Gemini, MiniMax). Qwen concurs directionally: new
accounts receive Tier 1 approval permitting only long calls and puts,
and spread strategies "are unlikely to be approved for an account with
only $100 in equity." MiniMax additionally claims a
$10,000+ floor for naked writing and spreads at Level
3/4; single-sourced, retained as the upper bound at [T5].
At USD 100, short premium is inaccessible on capital grounds
regardless of which figure is correct. MiniMax's own
falsification appendix makes the point sharply: at USD 100 the maximum
position is one cash-secured put on a $1-strike underlying — which does
not exist in the liquid universe.
Verdict: the VRP is the most economically robust effect in this section and is structurally unavailable at USD 100. That combination — real, harvestable, and gated behind a 20× capital requirement — is the section's central irony.
Evidence. Mitchell & Pulvino (2001),
"Characteristics of Risk and Return in Risk Arbitrage," Journal of
Finance 56(6), 2135–2175, DOI 10.1111/0022-1082.00418
[T1] (Source: Gemini). This citation replaces
MiniMax's, whose sole [T1] merger-arbitrage source could
not be located and appears fabricated, and whose supporting citations
included a paper about venture-capital valuation waterfalls with no
bearing on merger spreads (Source: MiniMax digest flags F12,
M7).
Effect size. Gemini: +3% to +6%
annualized with moderate decay. MiniMax: 1–5% annualized risk
premium above the risk-free rate, narrowing as event-driven hedge funds
crowded in. Overlapping ranges; adopted: 2–6%
annualized [T2]. At the individual-deal level
MiniMax gives the more usable decomposition: an announced deal typically
trades at a 5–10% spread to the offer price at announcement, narrowing
to 1–3% over the final 30–60 days for low-risk cash
deals with regulatory approval in hand [T4].
Risk profile. Both quantified sources agree the return is a short-put payoff on deal completion: MiniMax states a broken deal loses 30–50% in a day; Gemini calls it "single-deal jump-to-default risk." At USD 100 there is no deal diversification — the position is a single binary with an asymmetric payoff of roughly +2% versus −40%. Reaching the target requires being right about a deal breaking, which is the short side and is not what the risk-arbitrage literature documents.
Minimum capital. Gemini: $1,000 minimum, $500 friction breakeven. MiniMax's Table C also says $1,000+, while its prose says USD 100 funds a single micro-lot — unreconciled inside that document (Source: MiniMax digest flag N8). On fractional-share platforms the micro-lot claim is mechanically true; the $1,000 figure is better read as the capital needed for the strategy (multiple deals) rather than a position. Both grade it not feasible for the objective.
Evidence. Harris & Gurel (1986), JF
41(4), 815–829, DOI 10.1111/j.1540-6261.1986.tb04550.x
[T1]; Shleifer (1986), "Do Demand Curves for Stocks Slope
Down?", JF 41(3), 579–590, DOI
10.1111/j.1540-6261.1986.tb04518.x [T1];
Wurgler & Zhuravskaya (2002), Journal of Business 75(4),
583–608, DOI 10.1086/341638 [T1] (Source:
MiniMax, all three DOIs verified). Gemini adds Madhavan (2003),
Financial Analysts Journal 59(4), 51–64 [T2] — the
DOI it supplies carries a Journal of Portfolio Management
prefix and is not used here.
Effect size. MiniMax supplies four mutually
incompatible magnitudes for the same effect within one document — 5–7bps
on announcement plus 15–25bps around the effective date (≈0.2–0.3%
total), then "2–4% in the 1–2 weeks surrounding the announcement," then
"shrunk to 1–2% total," then "~2–4% per event" in its table (Source:
MiniMax digest flag N3). That internal incoherence disqualifies
MiniMax's figures as stated. Gemini gives +1.5% to +3.0% per
rebalance with high institutional decay. MiniMax's own
post-decay figure ("1–2% total") overlaps Gemini's lower bound.
Adopted: 1.5% to 3.0% per event, decayed substantially from the
1986-era effect by ETF-driven arbitrage [T2].
Feasibility. Executable at USD 100 as a single-name
binary and worth roughly 0.5–1.5% net of costs on MiniMax's estimate
[T6]. Gemini contributes the decisive constraint the other
two miss: the Russell reconstitution is annual, in June
[T5]. A 90-day window either contains a reconstitution
event or it does not, and if it does not, the strategy has zero trading
opportunities. Graded No on expected-return grounds by
both quantified sources, and additionally on calendar grounds by
one.
This subsection carries the most important epistemic distinction in the section, and it must not be blurred: there is a substantial peer-reviewed literature on prediction-market efficiency and on betting-market biases, and there is essentially no peer-reviewed literature on retail profitability in CFTC-regulated event-contract venues. Every claim below is labeled as DIRECT or TRANSFERRED accordingly.
Prediction markets aggregate dispersed information and produce
probabilities that frequently outperform individual experts and simple
statistical models [T1]; forecast accuracy improves as
events approach resolution (Source: Qwen, MiniMax). Manski
(2006), "Interpreting the Predictions of Prediction Markets,"
Economics Letters 91(3), 425–429, DOI
10.1016/j.econlet.2005.10.008 [T1], is the
standard reference on the gap between contract prices and mean beliefs.
Berg, Nelson & Rietz (2008), "Prediction Market Accuracy in the Long
Run," International Journal of Forecasting 24(2), 285–300, DOI
10.1016/j.ijforecast.2008.03.007 [T1],
establishes long-run accuracy on the Iowa Electronic Markets. Brier
(1950), Monthly Weather Review 78(1), 1–3, supplies the
standard scoring rule [T1]. (Source: MiniMax §13.1,
which its own verification pass identifies as the single most reliable
subsection in that file.)
MiniMax reports Brier scores of 0.02–0.05 on mature markets (IEM,
Betfair) [T6] — uncited, implausibly low for a 0–1 scale,
and presented with false precision. Not adopted. The
qualitative claim that mature political markets are well calibrated
stands [T1]; the numeric claim does not.
The favorite-longshot bias (FLB) is the best-documented exploitable
regularity in betting markets: low-probability outcomes are
systematically overpriced and high-probability outcomes systematically
underpriced [T1]. Canonical sources: Sauer (1998), "The
Economics of Wagering Markets," Journal of Economic Literature
36(4), 2021–2064 [T1]; Snowberg & Wolfers (2010),
"Explaining the Favorite-Longshot Bias: Is it Risk-Love or
Misperceptions?", Journal of Political Economy 118(4), 723–746
[T1], which attributes the bias to Prospect-Theory
probability weighting π(p) > p rather than to
risk-loving preferences. (Source: Gemini, MiniMax — the two digests
supply conflicting DOIs for Snowberg & Wolfers,
10.1086/655443 vs 10.1086/655844; the conflict
is unresolved and neither DOI is asserted here.) Political-market
FLB is documented in Rhode & Strumpf (2004), Journal of Economic
Perspectives 18(2), 127–141 [T1] and Berg & Rietz
(2003), Information Systems Frontiers 5(1), 79–93
[T1].
Magnitude — TRANSFERRED, racetrack. Longshots
overpriced by ~25–30%; favorites underpriced by ~3–5% [T1],
attributed to Snowberg & Wolfers (2010) (Source: MiniMax).
Political markets show a smaller, direction-dependent bias — more
efficient than racetracks, with non-trivial miscalibration remaining
only in the low-probability tail [T2].
Magnitude — claimed DIRECT, rejected. Gemini asserts
that buying underpriced favorites at P ≥ 0.70 on
Kalshi/ForecastEx yields +5.0% to +12.0% gross EV per
contract, refined in its executive summary to +6.4% per
trade net of maker fees, and on this basis ranks the strategy
#1 with P(reach $200) = 0.18 (Source: Gemini Table C,
Table G, JSON strategies[0]). This is
downgraded to [T6] and is not carried as a
finding. Reasons: (a) it is single-sourced; (b) Gemini's own
audit states that the Table C effect-size ranges "do not appear in those
papers as stated; they are the report's own harmonizations"; (c) it is
two to four times the directly-sourced racetrack favorite-underpricing
figure of 3–5%; (d) the derivation from a range to the 6.4% point
estimate is never shown; (e) its net-of-fee status depends on the
"Kalshi maker orders = 0% fee" assumption, which Gemini asserts with no
fee-schedule citation while carefully citing the taker formula
— its own audit flags this as load-bearing and unsupported. The
underlying phenomenon (FLB exists and favors buying favorites) is
[T1]; the tradeable magnitude on a US-regulated retail
venue is [T6].
MiniMax's competing magnitudes — a Polymarket FLB of "≈5–10%" and retail informed-trader returns of "1–5% per trade" — are struck entirely. Both are sourced solely to "Penn, C. (2025), 'An Empirical Study of Prediction Markets,' forthcoming International Journal of Forecasting," a paper that does not exist. A forthcoming IJF paper with a claimed March 2025 preprint would be findable; its absence is strong evidence of fabrication, and the surname coincidence with this project's commissioner makes the fabrication mechanism legible (Source: MiniMax digest flag F1).
Fee drag — the constraint that survives all of this.
Kalshi's taker fee is ceil(0.07 × P × (1 − P) × N) / 100,
peaking at roughly $0.02 per contract per side near P = 0.50 and falling
toward zero at the price extremes [T5] (Source: Qwen,
Gemini — two-source agreement on the formula). The ceiling function
is what matters at micro-notional: Gemini computes it as a 10%
drag on a $0.10 bet [T5]. MiniMax states
independently that Kalshi fees "scale non-linearly with price and wash
out small-notional edges." Three of three sources agree that the
fee structure is regressive against small stakes. Note the
interaction with the FLB: the bias is largest at the price extremes,
where the fee is smallest — a favorable alignment — but the return
per contract on a favorite at P = 0.70 is capped at
(1 − P)/P ≈ 43% if the favorite wins, against a 30% chance
of losing 100% of the stake. Repeated favorite-buying therefore has
positive expected return only if the true probability exceeds the price
by enough to clear fees. Author derivation
[T6]: at 1.43× per win, doubling requires two consecutive
full-stake wins (1.43² ≈ 2.04), which occurs with
probability 0.70² = 49% in a fair market —
attractive-looking until it is read correctly, because it is simply the
observation that a 49%-likely double against a 51%-likely ruin is close
to a coin flip with no edge, which is what a correctly priced contract
should deliver. Any excess over 49% must come entirely from the FLB
edge, and the transferred racetrack estimate of that edge is 3–5%, not
the 5–12% Gemini claims.
MiniMax states the position with a clarity that should be preserved
verbatim in the master report: "Direct peer-reviewed studies of
retail-account profitability on Kalshi, ForecastEx, or Polymarket are
practically nonexistent as of 2026-08-01," and "Any positive
claim about retail profitability in these venues is unsupported by the
peer-reviewed literature" [T6] for the retail-specific
inference, [T1] for the absence claim itself insofar as it
is a statement about the corpus (Source: MiniMax §13.2). This
is the most valuable single sentence in MiniMax's cluster and it
survives the citation audit intact because it is a negative claim
requiring no citation to support.
The non-peer-reviewed signals that do exist point the same direction and are graded accordingly:
[T4], single-sourced to a sports-media
explainer (Source: Qwen). The "up to" hedge makes the claim
unfalsifiable; retained as directional colour only.[T4], attributed by Qwen to "a study of
Kalshi" while its two supporting sources — a Yale SOM Insights piece and
a social-media post about an SSRN paper — both concern
Polymarket (Source: Qwen digest §7e). The
platform in the claim does not match the platform in the sources.
Retained at [T4] with the mismatch flagged; not
upgraded, not treated as a Kalshi finding. The directional
content — that gains concentrate in a small cohort of informed traders —
is consistent with the racetrack and political-market literatures and
with MiniMax's independent conclusion that "institutional traders
dominate documented positive-EV opportunities."MiniMax's cluster does one thing better than either sibling report: it states the transfer assumption and then attacks it. Reproduced because the master report needs it (Source: MiniMax §13.2):
Findings from professional and institutional prediction-market traders and from racetrack bettors carry over to retail event contracts only under the assumption that the markets are sufficiently homogeneous — the same FLB, the same calibration biases, the same information-asymmetry mechanisms operating at retail scale. This assumption is contestable. Retail traders may be less informed, less skilled, and more easily selected against by liquidity providers than the institutional traders in the source literature.
| Source literature (DIRECT) | Retail translation (TRANSFERRED) | Assumption required for the transfer to hold |
|---|---|---|
Snowberg & Wolfers (2010) — racetrack FLB [T1] |
Retail faces FLB on event markets; longshots overpriced | The same probability-misperception mechanism operates at retail scale and in economic/political categories |
Berg & Rietz (2003) — IEM calibration [T1] |
Retail calibration in political markets comparable to institutional | Venues are efficient enough that retail observes the same residual miscalibration |
Wolfers & Zitzewitz — informed-trader edge [T1],
identifier disputed |
A well-informed retail trader can extract positive EV on thin contracts | The retail trader has information comparable to the institutional informed trader |
The four documented conditions under which positive expected value is
attainable, as MiniMax reports them [T2]: (1) persistent
miscalibration; (2) categories with structural information asymmetry;
(3) markets illiquid enough that a single position moves the price; (4)
high capital and rapid execution. Conditions (3) and (4) are
mutually hostile at USD 100 — a stake that can move a thin
market is a stake that cannot exit it, and condition (4) explicitly
excludes the subject of this report.
Verdict. The probability of reaching 2× in 90 days
via prediction markets at USD 100 is low and is not supported by
the peer-reviewed literature [T1] for the absence,
[T6] for any specific probability. Gemini's
P(reach $200) = 0.18 for this strategy is the highest such
figure anywhere in the merged corpus and it is uncited, unmodeled, and
built on a fee assumption its own author flagged.
One further conflict is deferred, not resolved here.
Qwen reports that Massachusetts regulators secured preliminary
injunctions against Kalshi, that a Suffolk County Superior Court judge
barred Kalshi from offering sports contracts in the state, and that an
MA resident "cannot currently participate in these federally regulated
markets without facing potential legal jeopardy" [T4].
Gemini reports the opposite: that CEA § 5c(c)(5)(C) and 17 CFR 40.11
preempt state gaming law on designated contract markets, that the
Massachusetts Gaming Commission's authority "does NOT apply to CFTC DCM
event contracts on economics/elections," and that Kalshi and ForecastEx
are therefore accessible [T2]. These are flatly
irreconcilable and the disagreement is decisive for the feasibility
column. It is a regulatory question, not a strategy-evidence
question, and is referred to the regulatory section of the master
report. Table C marks the prediction-market rows' feasibility as legally
contingent.
The VRP implies a structural asymmetry: the long-premium
buyer pays it and the short-premium seller earns it. MiniMax
puts long-volatility expected return on broad-index options at roughly
−5% to −10% annualized [T6] — the figure
rests on a citation that could not be located, so it is not adopted as a
measurement, though its sign is corroborated three ways. Coval &
Shumway (2001) [T1] is the correct anchor for the sign:
zero-beta straddles earn large negative returns (Source: MiniMax
digest flag M5, verification note).
Qwen frames the tension usefully and reaches the opposite operational conclusion from the other two: selling premium produces "a steady stream of income (theta decay)" but the maximum gain is capped at the premium collected, so it cannot double the account — while buying premium has a payoff profile (loss capped at premium, gain theoretically unlimited) that matches the first-passage requirement of a large positive jump (Source: Qwen). This is analytically correct and worth preserving: under a fixed-multiple, fixed-deadline objective, the strategy with positive expected value cannot reach the target and the strategy that can reach the target has negative expected value. Qwen's own conclusion — that buying deep OTM options is "a lottery ticket" — closes the loop. All three reports agree on the endpoint; only Qwen states the structural reason cleanly.
MiniMax identifies the counterparties [T4]:
market-makers and dealers, banks and structured-product issuers writing
autocallables and return-enhancement notes, volatility-selling hedge
funds, and institutional index put-writers including pensions, insurers,
and buy-write programs. Its structural claim: "Historically, the
retail side is the buyer of vol, not the seller." Single-sourced
and uncited at the level of specific firms [T6], but the
directional claim is corroborated by Gemini's finding that retail
long-premium directional trades carry negative EV, and by the
account-minimum evidence in §5.4.4 — retail cannot be the seller
because retail cannot obtain the approval tier or post the
margin. The exclusion is mechanical, not behavioural.
Only MiniMax treats 0DTE mechanics; Gemini treats it only as a negative-EV vehicle; Qwen does not mention it. Single-sourced content is tagged accordingly.
[T5],
single-sourced.[T4], single-sourced.[T3], single-sourced,
working-paper stage.[T3],
single-sourced, working-paper stage. If true, this makes retail's
position worse rather than better, since retail is the buyer.Gemini grades 0DTE long directional at negative EV of −15% to
−30% per trade, with bid-ask spread drag of 10–30% of premium
and extreme theta decay, minimum $10, NOT feasible
(Source: Gemini Table C, JSON strategies[3]). The
magnitude is a Gemini harmonization [T6]; the sign and the
spread-drag mechanism are corroborated. MiniMax's verdict is identical
in substance: 0DTE binary bets are lottery trading with negative
expected return, large variance, and a median outcome of total
loss. A 0DTE put spread does scale to USD 100 notionally — and
is on the wrong side of the VRP.
Three reports agree on direction; the magnitudes conflict irreconcilably and most are discarded.
Best-supported citation. Bryzgalova, Pavlova &
Sikorskaya (2023), "Retail Trading in Options and the Rise of the Big
Three Wholesalers," Journal of Finance 78(6) [T1].
Both Gemini and MiniMax cite this work, each with a defective author
list or venue (Gemini as a 2023 MIT IDE working paper; MiniMax as
"Bryzgalova, Pavlova (2024), Journal of Finance, working paper"
— a category error, since a journal does not host working papers). The
corrected form above is the one supported by MiniMax's verification
pass, and the paper's existence is corroborated two ways (Source:
Gemini, MiniMax).
Additional canonical source, surfaced by audit rather than by
any report body. Barber, Huang, Odean & Schwarz (2022),
"Attention-Induced Trading and Returns: Evidence from Robinhood Users,"
Journal of Finance 77(6), 3141–3190 [T1].
MiniMax's verification pass names this as the canonical peer-reviewed
retail-performance paper conspicuously absent from a section entirely
about retail performance (Source: MiniMax digest §5.6). It is
included here as a gap-fill with its provenance labeled: it was
identified by the audit, not asserted by any of the three reports.
Magnitudes, graded individually:
[T6]. Single-sourced, attributed
to a practitioner journal that could not be located in a single
verification pass. The venue (Journal of Risk Management in
Financial Institutions) is poorly indexed, so a miss is weak
evidence of fabrication — but the figure is nonetheless uncorroborated
and is not carried as a finding.[T6]. Struck. This is
arithmetically irreconcilable with a sub-10% win rate on capped-loss
instruments, as MiniMax's own audit notes, and MiniMax's Table C
propagates it anyway (Source: MiniMax digest flag N1).[T4]. Attributed only to unnamed
industry data — a headline number with no identifiable source.
Retained at [T4] with the sourcing defect stated;
not upgraded.[T6], Gemini's harmonization, sign corroborated, magnitude
not.What survives. The direction is agreed three ways
and is the only claim worth carrying at strength: retail
long-option trading has documented negative expected return, and the
peer-reviewed evidence base for it is narrower than any of the three
reports implies — limited to specific markets and specific time
periods, with most post-2020 work still at working-paper stage
[T3]. MiniMax's summary is the right one and does not
depend on any disputed number: "The retail options trade is one of
the few trading strategies at USD 100 that produces a documented
negative expected return on every trade." Restated more carefully
to match the evidence: the expected return is negative in the population
average, on capped-loss instruments, before considering that the median
outcome of a short-dated OTM long position is total loss of the
premium.
Rows merged across all three sources. Where sources conflict, the resolved value appears and the conflict is recorded in Resolved Conflicts. Effect sizes are gross unless stated. "Feasible at $100/90d" answers only whether the strategy can be executed and can plausibly reach USD 200 — a "No" on capital grounds and a "No" on effect-size grounds are distinguished in the failure column of the prose above.
| Strategy | Evidence tier | Key citation (DOI) | Documented effect size | Post-publication decay | Min. viable capital | Friction breakeven | Feasible at $100 in 90d? |
|---|---|---|---|---|---|---|---|
| Market excess return (Mkt-RF) | [T1] |
Sharpe (1964), 10.1111/j.1540-6261.1964.tb02865.x; Fama
& French (1993), 10.1016/0304-405X(93)90023-5 |
~6–8% annualized real | Low; no literature argues it has disappeared | $1 (fractional ETF) | <0.1% (spread only) | No — horizon; ~1.7% expected over 90d |
| Size (SMB) | [T2] |
Fama & French (1993),
10.1016/0304-405X(93)90023-5 |
~2% annualized | Modest; contested post-1980 | $1,000+ (portfolio) | n/a at this scale | No — horizon-infeasible |
| Value (HML) | [T2] contested |
Fama & French (1993/2015),
10.1016/j.jfineco.2014.10.010 |
~0–3% annualized | Severe 2017–2020; partial post-2021 rebound; rejected as independent by q-model | $1,000+ | ~0.1% per rebalance | No — sub-1% over 90d |
| Momentum, cross-sectional 3–12m | [T1] effect, [T2] magnitude |
Jegadeesh & Titman (1993),
10.1111/j.1540-6261.1993.tb04702.x; Asness, Moskowitz &
Pedersen (2013), 10.1111/jofi.12021 |
4–8% annualized post-decay (~0.5%/mo) | ~50–58% (McLean-Pontiff) | ~$10/leg; ~$300 for a 10-name sleeve | 0.5–1% per rebalance budgeted; 4–8% actual at retail | No — friction exceeds budget 8–16×; 90d too short |
| Time-series momentum | [T1] |
Moskowitz, Ooi & Pedersen (2012),
10.1016/j.jfineco.2011.11.003 |
Comparable to cross-sectional; long-only | Not separately quantified in corpus | $10 (single fractional) | 0.5% per rebalance | No — same horizon and friction failure |
| Short-term reversal (1-week) | [T1] effect, [T2] magnitude |
Jegadeesh (1990), 10.1111/j.1540-6261.1990.tb05110.x;
Lehmann (1990), 10.2307/2330889 |
0.5–1.0% per week gross | Substantial; compressed by electronic market-making | $50–100 (single name) | $500 equivalent; 80%+ of gross consumed by fees | No — net over 90d straddles zero (−5.5% to +9%) |
| Profitability (RMW) | [T1] |
Novy-Marx (2013), 10.1016/j.jfineco.2013.01.003 |
~0.5%/month [T6] on magnitude |
Modest | $300 (portfolio) | ~0.3% per rebalance | No — ~1.5% over 90d |
| Investment (CMA) | [T1] |
Cooper, Gulen & Schill (2008),
10.1111/j.1540-6261.2008.01369.x; Titman, Wei & Xie
(2004), 10.1017/S0022109000003125 |
~0.3%/month [T6] on magnitude |
Modest | $300 | ~0.3% per rebalance | No |
| Quality (QMJ) | [T1] effect, [T2] independence |
Asness, Frazzini & Pedersen (2019),
10.1007/s11142-018-9470-2 |
~0.4%/month [T6] on magnitude |
Modest | $300 | ~0.3% per rebalance | No |
| Betting Against Beta (BAB) | [T1]; subsumption claim [T6] |
Frazzini & Pedersen (2014),
10.1016/j.jfineco.2013.10.005 |
~0.5%/month [T6]; premium peaks in financial
stress |
Contested — subsumption by standard risk factors claimed but unverified | $1,000+ and margin | 0.5% per rebalance | No — requires shorting; margin account unavailable |
| Idiosyncratic / low volatility | [T2] |
Sole citation could not be verified; effect corroborated only by AQR
grey literature [T4] |
~0.4%/month [T6] |
Modest; material trading-cost deduction | $300 | ~0.5% per rebalance | No |
| Carry (FX / bond / commodity) | [T1] |
Koijen, Moskowitz, Pedersen & Vrugt (2018),
10.1016/j.jfineco.2017.11.002 |
4–8% annualized | Modest; survives across asset classes | $10,000+ (futures account) | n/a at this scale | No — institutional infrastructure required |
| PEAD | [T1] effect, [T2] magnitude |
Ball & Brown (1968), 10.2307/2490232; Bernard &
Thomas (1989), 10.2307/2491256; Bernard & Thomas
(1990), 10.1016/0165-4101(90)90008-R |
+2% to +5% on long leg over 30-day hold, top-decile SUE | ~35–50%; decay citation in one source does not support the claim | $5 (fractional); ~$300 for 10-event diversification | ~0.3% per round trip; needs sub-5bp spreads | Marginal on execution, No on target — ≤ ~16% compounded over 3 events |
| Long-term reversal (3–5y) | [T1] |
DeBondt & Thaler (1985),
10.1111/j.1540-6261.1985.tb05004.x |
~5% annualized | Significant | $1,000+ | 0.5% per rebalance | No — horizon exceeds mandate by 12–20× |
| Accruals | [T2] |
Sloan (1996), Accounting Review 71(3), 289–315 (DOI disputed across sources) | ~2–4% annualized, materially reduced | Substantial; partly subsumed by profitability | $1,000+ | ~0.3% per rebalance | No |
| Net stock issuance | [T2] |
Loughran & Ritter (1995), JF 50(1) (DOI internally inconsistent in source; not asserted) | ~2–4% annualized | Substantial | $1,000+ | ~0.3% per rebalance | No |
| Volatility risk premium — short premium | [T1] effect, [T2] magnitude |
Carr & Wu (2009), 10.1093/rfs/hhn038; Coval &
Shumway (2001), 10.1111/0022-1082.00352 |
+1% to +2% per month [T2], single-sourced;
crash risk: −30% to −90% in a single session (Feb 2018, Mar
2020) [T4] |
Narrowed post-2014; still positive | $2,000 (cash-secured put margin); $10,000+ for
naked/spread [T5] |
0.5–2% per round trip on liquid SPX/SPY | No — capital and approval-tier gated 20× above stake |
| Long premium / long volatility | [T1] sign |
Coval & Shumway (2001),
10.1111/0022-1082.00352 |
Negative EV; magnitude of −5% to −10% annualized is
[T6], uncorroborated |
n/a | $5–100 (one contract) | ~0.5% per round trip plus 10–30% spread cross | No — structurally on the wrong side of the VRP |
| 0DTE long directional (retail) | [T3]/[T4]; magnitude
[T6] |
Bryzgalova, Pavlova & Sikorskaya (2023), JF 78(6); Cboe
volume data [T4]; SSRN working papers 2023–2025
[T3] |
Negative EV; −15% to −30% per trade [T6]; retail loses
65–80% of premium over 12-month windows [T4], unnamed
source |
n/a — market is post-2022 | $10 (one cheap contract) | Bid-ask 10–30% of premium; theta drag extreme | No — median outcome is total loss of premium |
| Retail options trading, general | [T1] direction, [T2] magnitude |
Bryzgalova, Pavlova & Sikorskaya (2023), JF 78(6); Barber, Huang, Odean & Schwarz (2022), JF 77(6), 3141–3190 (DOI not supplied by any source; not asserted)* | Negative EV in population average; "<10% of trades profitable" is
[T6], single-sourced and unverified |
n/a | $100 (one contract) | Variable; spread plus theta | No — structurally lossy |
| Merger arbitrage / event-driven | [T1] |
Mitchell & Pulvino (2001),
10.1111/0022-1082.00418 |
2–6% annualized; 1–3% per low-risk deal over 30–60 days; broken deal −30% to −50% in one day | Narrowed as event-driven funds crowded in | $1,000 (strategy); $100 funds one micro-lot position | ~0.5% per round trip; $500 breakeven | No — single-deal binary; payoff +2% vs −40% |
| Index reconstitution | [T1] effect, [T2] magnitude |
Harris & Gurel (1986),
10.1111/j.1540-6261.1986.tb04550.x; Shleifer (1986),
10.1111/j.1540-6261.1986.tb04518.x; Wurgler &
Zhuravskaya (2002), 10.1086/341638 |
1.5–3.0% per event, decayed from the 1986-era effect | Substantial — ETF-driven arbitrage | $100 (single event) | ~0.3–1.5% per round trip; net alpha ~0.5–1.5% | No — expected return insufficient; Russell reconstitution is annual (June), so a 90-day window may contain zero events |
| Prediction market — favorite buying (FLB harvest) | [T1] phenomenon; [T6] tradeable
magnitude |
Snowberg & Wolfers (2010), JPE 118(4), 723–746 (DOI disputed across sources; not asserted); Sauer (1998), JEL 36(4) | TRANSFERRED (racetrack): longshots overpriced
~25–30%, favorites underpriced ~3–5% [T1]. Claimed
direct (Kalshi, P≥0.70): +5% to +12% EV —
[T6], rejected as uncorroborated
harmonization |
Persistent in retail-dominated venues [T2]; no decay
series exists for regulated event contracts |
$1.00 (single contract) | Kalshi taker fee ceil(0.07·P·(1−P)·N)/100, ~10% drag on
a $0.10 bet [T5]; "maker = 0%" is [T6],
uncited |
Legally contingent — deferred. Two sources
contradict each other on MA accessibility. On economics alone:
No — at P=0.70 doubling needs 2
consecutive full-stake wins, ≈49% under a fair
market against ≈51% ruin [T6] author
derivation; the edge over that coin flip is the 3–5% FLB, less fees |
| Prediction market — informed / asymmetric-information trading | [T2] |
Wolfers & Zitzewitz, JEP (year, volume, and DOI disputed across sources; not asserted) | Positive EV documented for well-informed traders; four enabling conditions, two of which exclude a $100 account | n/a | $100 (single contract) | Kalshi fee formula as above | No — documented edge accrues to high-capital,
fast-execution informed traders; retail-specific evidence is nonexistent
[T1] for the absence |
| Prediction market — retail profitability, generally | [T1] for the absence of evidence |
No peer-reviewed study exists as of 2026-08-01 quantifying retail Sharpe or hit rates on Kalshi, ForecastEx, or Polymarket | No documented effect size exists. Grey-lit signals:
"up to 80% of users are net losers" [T4]; "top 1% capture
84% of gains" [T4], platform mismatch flagged |
n/a | n/a | n/a | No — any positive claim is unsupported by the peer-reviewed literature |
* Barber, Huang, Odean & Schwarz (2022) was surfaced by MiniMax's verification pass as a canonical omission, not asserted by any of the three report bodies. Journal, volume, and page range are as supplied by that audit; no DOI is asserted, because none appears anywhere in the corpus and none is invented here.
Row count: 25 strategies. Duplicate rows present in MiniMax's original Table C (momentum and short-term reversal each appeared twice with different effect sizes) have been merged.
Anomaly replication-failure rate. MiniMax:
"roughly 80–90% of published anomalies fail," uncited. Gemini: 65%+ fail
at t ≥ 1.96, 82%+ at t ≥ 2.78, cited to
Hou-Xue-Zhang (2020) with a correct DOI. Qwen: no figure.
Resolved to Gemini — better-cited,
threshold-conditional, and MiniMax's own verification pass independently
identifies 65% as the paper's verified headline and flags its 80–90% as
unsupported. MiniMax's figure pruned.
Harvey-Liu-Zhu venue and DOI. MiniMax:
Review of Finance 21(1), 1–33, DOI
10.1093/rof/rfv003. Gemini: Review of Financial
Studies 29(1), 5–68, DOI 10.1093/rfs/hhv059.
Resolved to Gemini; MiniMax's own audit confirms its
journal and DOI are both wrong. Factor count reported as "roughly 315"
rather than MiniMax's flat 316, since the paper's count is 313 or 316
depending on the counting convention.
McLean-Pontiff DOI and decay decomposition. Both
quantified sources agree on 58% total decay. MiniMax's DOI
(10.1111/jofi.12349) is wrong per its own audit;
Gemini's 10.1111/jofi.12365 adopted. On
decomposition, MiniMax says "roughly half over-fitting, remainder
publication"; Gemini gives 26% over-fitting + 32% arbitrage.
Gemini adopted — more specific and consistent with
MiniMax's qualitative statement.
Existence of a replication crisis. MiniMax's
framing (80–90% failure) is directly contradicted by Jensen, Kelly &
Pedersen (2023), which MiniMax cites approvingly while suppressing its
headline conclusion that most factors do replicate. Neither Gemini nor
Qwen mention the paper. Not resolved by majority — both branches
presented explicitly, with the observation that the objective's
conclusion is invariant to which branch is correct. The JKP DOI is
disputed (10.1111/jofi.13255 given,
10.1111/jofi.13249 proposed as the correction); no
DOI asserted.
PEAD effect size. MiniMax: 7–10% over 60 days (original), 3–5% modern. Gemini: +2–5% over a 30-day hold, 35% decay. Qwen: qualitative only, and attributes PEAD to Loughran & Ritter (1995) — a paper about IPO underperformance. Qwen's attribution discarded. MiniMax and Gemini unified on horizon-adjustment to +2% to +5% over 30 days; MiniMax's 7–10% original figure is flagged by its own audit as likely inflated against Bernard & Thomas's canonical hedge return and is not carried forward.
PEAD citation DOI. Gemini gives Bernard &
Thomas (1989) as JAR 27 with an Elsevier JAE DOI
prefix, a mismatch its own audit flags. MiniMax gives JAR 27
Supplement, 1–36, DOI 10.2307/2491256, verified in its own
audit. MiniMax adopted on citation quality.
PEAD decay citation. MiniMax's ~50% decay figure is sourced to a paper about profitability and book-to-market that does not study PEAD (flag F20/M4). The citation is struck; the decay claim survives only at Gemini's uncited 35%, so the decay is reported as directionally large and imprecisely measured rather than as a number.
Short-term reversal effect size and feasibility. MiniMax: ~0.5%/week, graded "Marginal." Gemini: 0.5–1.0%/week, 80%+ consumed by fees, $500 breakeven, graded "NO." Qwen: qualitative, misattributed to Jegadeesh & Titman (1993). Effect size unified at 0.5–1.0%/week; feasibility resolved to "No" — MiniMax's own prose (4–12% friction over 90 days, "eating most of the documented effect size") contradicts its own table grade, and Gemini agrees with MiniMax's prose. Qwen's attribution discarded.
Momentum citation. Gemini attributes cross-sectional momentum to Harvey, Liu & Zhu (2016), a multiple-testing critique — a misattribution its own audit identifies. Replaced with Jegadeesh & Titman (1993) and Asness, Moskowitz & Pedersen (2013) from MiniMax, both verified.
Momentum effect size and feasibility. MiniMax: ~0.5%/month post-decay, Table C grade "Marginal." Gemini: 4–8% annualized, grade "NO." Qwen: infeasible. Effect sizes are compatible (0.5%/mo ≈ 6%/yr) and unified at 4–8% annualized. Feasibility resolved to "No" 3–0, overriding MiniMax's table grade using MiniMax's own friction arithmetic (4–8% per round trip against a 0.5–1% budget).
Volatility risk premium magnitude. MiniMax gives
~0.10%/day in prose and ~0.5%/day in its table — a 5× internal
contradiction — and attributes a short-straddle Sharpe of 0.50–0.75 to
Coval & Shumway (2001), a figure its verification pass found
invented (the paper reports large negative straddle returns).
Both MiniMax magnitudes and the 0.50–0.75 Sharpe are
struck. Gemini's +1–2%/month adopted at [T2],
single-sourced and explicitly labeled as a report
harmonization.
Carr & Wu (2009) DOI. Gemini supplies a
Journal of Banking & Finance prefix for an RFS
article, flagged by its own audit. MiniMax supplies
10.1093/rfs/hhn038, verified. MiniMax
adopted.
Short-premium minimum capital. MiniMax: $2,000
for cash-secured puts, $10,000+ for spreads and naked writing at Level
3/4. Gemini: $2,000. Qwen: Tier 1 only for new accounts, spreads need
higher tiers and "are unlikely to be approved for an account with only
$100." $2,000 adopted as the binding floor (2 of 3 explicit,
third concurs directionally); MiniMax's $10,000 retained as the upper
bound for naked/spread writing at [T5],
single-sourced. The verdict is unaffected either way — both
figures are 20× to 100× the stake.
Merger-arbitrage citation. MiniMax's sole
[T1] source could not be located and appears fabricated;
its supporting citations include a venture-capital valuation paper with
no bearing on merger spreads. Gemini supplies Mitchell & Pulvino
(2001), JF 56(6), 2135–2175, DOI
10.1111/0022-1082.00418 — the canonical paper, which
MiniMax's own audit names as the conspicuous omission. Gemini's
citation adopted; MiniMax's fabricated citation dropped
entirely.
Merger-arbitrage effect size and minimum capital. MiniMax: 1–5% annualized, Table C 2–5%, min capital $1,000+ contradicting its own prose claim that $100 funds a micro-lot. Gemini: 3–6% annualized, min $1,000. Unified to 2–6% annualized. Minimum capital resolved as $1,000 for the strategy and ~$100 for a single position — the two figures answer different questions and both are retained with that distinction stated.
Index-reconstitution effect size. MiniMax supplies four mutually incompatible magnitudes in one document (0.2–0.3%, 2–4%, 1–2%, 2–4%). Gemini gives 1.5–3.0% per rebalance. MiniMax's figures disqualified as internally incoherent; Gemini's 1.5–3.0% adopted, noting that MiniMax's own post-decay figure (1–2%) overlaps it. Gemini's Madhavan (2003) DOI carries a JPM prefix for an FAJ article and is not asserted; the verified Harris & Gurel / Shleifer / Wurgler & Zhuravskaya citations from MiniMax are used instead.
Favorite-longshot bias magnitude — the section's most
consequential conflict. Gemini: +5% to +12% gross EV on Kalshi
favorites at P ≥ 0.70, refined to +6.4% per trade net,
feasibility YES, ranked #1 with
P(reach $200) = 0.18. MiniMax: racetrack favorites
underpriced 3–5%, longshots overpriced 25–30%, retail edge unsupported.
Qwen: FLB makes longshot betting a losing proposition; top cohort
captures the gains. Resolved against Gemini. Its figure
is single-sourced, is 2–4× the directly-sourced racetrack figure, is
described by Gemini's own audit as a report harmonization that does not
appear in the cited paper, and depends on an uncited "maker fee = 0%"
assumption that its own audit flags as load-bearing. Downgraded
to [T6] and excluded as a finding; the underlying
phenomenon retained at [T1] with the racetrack magnitudes
labeled TRANSFERRED.
Snowberg & Wolfers DOI. MiniMax gives
10.1086/655443; Gemini gives 10.1086/655844.
Both digests' verification passes endorsed their own version.
Irreconcilable on available evidence — neither DOI
asserted. The paper is cited by author, year, journal, volume,
and pages only.
Wolfers & Zitzewitz identifier. MiniMax
cites a 2006 JEP 20(2), 107–126 article with DOI
10.1257/jep.20.2.107; Gemini cites a 2004 JEP
18(2), 107–126 article with DOI 10.1257/0895330041371321;
MiniMax's own audit separately refers to a 2004 paper under a third
title. Irreconcilable — no year, volume, or DOI
asserted. The substantive claim (positive EV is documented for
well-informed prediction-market traders) is corroborated by two sources
and retained at [T2].
Polymarket FLB magnitude and retail edge — struck as fabricated. MiniMax's "≈5–10% Polymarket FLB," "median position size on Polymarket is too small to arbitrage miscalibration," and "retail informed traders run 1–5% per trade" all rest solely on "Penn, C. (2025), 'An Empirical Study of Prediction Markets,' forthcoming International Journal of Forecasting" — a paper that does not exist. A forthcoming IJF article with a claimed March 2025 preprint would be indexed; its absence is strong evidence of fabrication, and the surname matches this project's commissioner. All four load-bearing claims struck; nothing sourced to this citation enters the merged section.
Prediction-market retail loss statistics. Qwen's
"up to 80% of users are net losers" (sports-media explainer) and "top 1%
captured 84% of all trading gains" (attributed to Kalshi but sourced to
two Polymarket references) are single-sourced grey literature
with a platform mismatch. Retained at [T4] with the defects
stated inline; not upgraded, not treated as corroboration for any
Kalshi-specific claim.
Brier scores of 0.02–0.05 on mature markets.
MiniMax, uncited, implausibly low for a 0–1 scale, false precision.
Dropped as a number; the qualitative claim that mature
political markets are well calibrated is retained at
[T1].
Kalshi maker fee. Gemini asserts "maker orders =
0% fee" with no fee-schedule citation while carefully citing the taker
formula, and its entire top-ranked strategy depends on it. The taker
formula ceil(0.07·P·(1−P)·N)/100 is corroborated by Qwen
and Gemini independently → [T5]. The maker-fee
claim is [T6], single-sourced and uncited, and the
feasibility verdict that rests on it is downgraded accordingly.
Massachusetts accessibility of regulated event contracts. Qwen: MA injunctions against Kalshi, Suffolk County Superior Court order, "cannot currently participate without facing potential legal jeopardy." Gemini: CEA § 5c(c)(5)(C) and 17 CFR 40.11 preempt state gaming law on DCMs; the MA Gaming Commission's authority does not reach economic and election contracts; Kalshi and ForecastEx feasible. Flatly irreconcilable and decisive for feasibility. Not resolved here — it is a regulatory question, referred to the master report's regulatory section. Table C marks the affected rows "legally contingent."
Retail options loss magnitude. MiniMax gives
three mutually incompatible figures: "<10% of trades profitable,"
"loses 0.5–1.5% of premium per trade," and "65–80% of premium lost over
12-month windows." Gemini gives −15% to −30% per trade for 0DTE.
The 0.5–1.5% figure is struck as arithmetically impossible
against a sub-10% win rate (MiniMax's own audit identifies this
and notes its Table C propagates the wrong figure anyway). The
<10% win rate is [T6] — sole source could not
be located, though its practitioner-journal venue is poorly indexed, so
it is labeled unverifiable rather than fabricated. The 65–80%
figure is [T4] — unnamed industry data.
Gemini's −15% to −30% is [T6] — its own
audit says the ranges are harmonizations. Only the direction
(negative EV) is carried at strength, corroborated 3 of
3.
Bryzgalova et al. citation form. Gemini cites a
2023 MIT IDE working paper; MiniMax cites "Bryzgalova, Pavlova (2024),
Journal of Finance, working paper" — a category error, since a
journal does not host working papers, and a recognized signature of
synthesized citations (MiniMax uses the construction six times).
Corrected form adopted: Bryzgalova, Pavlova & Sikorskaya
(2023), "Retail Trading in Options and the Rise of the Big Three
Wholesalers," Journal of Finance 78(6), per MiniMax's
verification pass. The paper's existence is corroborated two ways; the
tier is corrected from [T3] to [T1] since it
is published.
Quality (QMJ) citation. MiniMax cites "Asness,
Frazzini & Israel (2019), AQR Working Paper" at [T4].
Corrected to Asness, Frazzini & Pedersen (2019), Review
of Accounting Studies 24(1), DOI
10.1007/s11142-018-9470-2 per MiniMax's own audit
— wrong third author and wrong venue, and since the paper is
peer-reviewed the [T4] grade was also wrong.
Betting Against Beta citation. MiniMax cites
"Frazzini, Kabiller & Pedersen (2018), JFE 130(1), 15–38."
Corrected to Frazzini & Pedersen (2014), JFE
111(1), 1–25, DOI 10.1016/j.jfineco.2013.10.005 —
Kabiller co-authored Buffett's Alpha, not BAB, and MiniMax has
the two papers' author sets and years swapped in both directions. The
BAB-subsumption counterclaim rests on a single citation
that could not be verified and never appears in MiniMax's own source
list; retained at [T6], not as an established
finding.
Idiosyncratic-volatility citation — downgraded, not
laundered. MiniMax's sole citation for the low-volatility
anomaly could not be located with the authors given, and one named
author is a market-structure researcher rather than an
idiosyncratic-volatility researcher, suggesting confabulated authorship.
Neither Gemini nor Qwen covers the factor. The row is retained
at [T2] with no key citation asserted, because the
low-volatility anomaly is genuinely well-established in the wider
literature — but no citation from this corpus is trustworthy enough to
attach to it, and none is invented to fill the gap.
Tier-taxonomy incompatibility across sources — a merge
hazard, recorded not resolved. Gemini redefined the CASINO
taxonomy, grading primary regulatory statutes [T1] and
dropping the replication requirement from [T1] entirely;
153 of ~194 of its tier tags are [T1]. Qwen used T1–T4
without ever defining the taxonomy and never used T5 or T6, grading a
VoxEU column [T1] while grading the factor literature
[T2]/[T3]. All tiers in this section were
re-assigned against the CASINO specification from the underlying
evidence, not inherited from any source report. Tier counts are
therefore not comparable to any source report's
counts.
Citations downgraded or dropped for fabrication or
unverifiability (consolidated): the "Penn (2025)"
prediction-markets paper (dropped entirely, four
load-bearing claims struck); MiniMax's sole merger-arbitrage
[T1] citation (dropped, replaced with
Mitchell & Pulvino); MiniMax's foundational VRP citation and its
"Volatility-of-Volatility Risk" citation (dropped,
replaced with Carr & Wu and Coval & Shumway); MiniMax's PEAD
decay citation (dropped — real-paper-wrong-topic);
MiniMax's momentum net-Sharpe source (dropped,
unlocatable); MiniMax's index-reconstitution price-impact citation
(dropped, wrong authors); MiniMax's low-volatility
citation (dropped, confabulated authorship, row
retained without a citation); MiniMax's retail-options win-rate source
(downgraded to [T6], unverifiable rather
than fabricated — poorly indexed venue); MiniMax's momentum-robustness
dissent citation (downgraded to [T6], same
reason); Gemini's Madhavan (2003) DOI, Carr & Wu DOI, and Bernard
& Thomas DOI (dropped, publisher-prefix
mismatches); Gemini's Harvey-Liu-Zhu-as-momentum attribution
(dropped); Gemini's "+5–12% Kalshi favorite EV," "+6.4%
net," and "maker fee = 0%" (downgraded to
[T6]); Qwen's Loughran-Ritter-as-PEAD and
Jegadeesh-Titman-as-reversal attributions (dropped);
Qwen's Wolfers-Zitzewitz-via-VoxEU [T1] grade and
Carr-Madan-via-BIS [T1] grade (dropped —
neither URL supports the graded claim).
Content gaps in this cluster, unfilled and flagged for the master report: no source treats the transaction-cost-of-anomalies literature (Novy-Marx & Velikov's taxonomy of anomalies and their trading costs), which is a conspicuous omission given that friction is the binding constraint throughout this section; no source treats Chen & Zimmermann's publication-bias work, the standard modern counterweight in the factor-zoo debate; no source contains CFTC- or Kalshi-specific empirical literature, which may reflect that none exists rather than a research failure; and no source provides a constructive options-strategy treatment — spread construction and volatility-surface trades are absent from all three, so this section's options coverage is unavoidably one-sided.
This section inventories named, cited statistical and forecasting
techniques developed outside finance and maps each to the $100 → $200 /
90-day problem. It answers the question the CASINO prompt refused to let
be answered with "use machine learning": every technique below is named,
attributed to an originating publication, given a concrete transfer
mechanism, and given a stated transfer risk. Where a source report named
a technique but stated no transfer risk, that omission is marked
GAP rather than filled in. Three source reports contribute:
MiniMax (dedicated Cluster 4 deliverable, eight domains, ~78 technique
rows across per-domain tables plus a 44-row Table D), Gemini (eight
domains covered one-to-one in a single 8-row matrix), and Qwen (four of
eight domains, 6-row matrix). Inline provenance markers name which
reports support each block.
Tier convention applied in this section. Source
reports applied the T1–T6 taxonomy inconsistently and both digests flag
tier inflation — Gemini graded primary regulatory statutes and textbooks
[T1] after redefining the scale, and MiniMax graded a NOAA
technical procedures bulletin and ten monographs [T1].
Gemini's digest carries an explicit merge instruction not to aggregate
T1 counts across reports. Tags below are therefore re-derived here, not
inherited:
[T1] — peer-reviewed primary research, independently
replicated, or a theorem proved in a refereed venue[T2] — peer-reviewed primary research not independently
replicated; canonical scholarly monographs restating refereed
results[T3] — preprints, theses, working papers[T4] — trade books, practitioner literature, vendor
marketing[T5] — primary regulatory, exchange, institutional, or
package documentation; grey literature[T6] — inference constructed by a source report or by
this merge; unverified or unsourced numericsThe structural fact that governs all eight domains
[T6] — asserted convergently by all three reports, none of
which attaches a primary citation to it. Every domain below developed
its methods against an exogenous data-generating process. The
atmosphere does not read the forecast. Case counts do not respond to the
nowcast. A manufacturing line holds no opinion about the control chart.
Test items do not become harder because a psychometrician estimated
their difficulty. Financial markets are the sole application domain in
this inventory where the data-generating process is populated by agents
who profit by eliminating exactly the regularity the technique detects.
All three reports converge on this framing independently, which is the
strongest agreement in the entire section. Qwen states it most cleanly:
"Weather patterns evolve according to physical laws that are not
influenced by the forecast itself. Financial markets, in contrast, are
populated by rational agents who constantly seek to exploit any
predictable patterns" [T6] — uncited in Qwen (Source:
Qwen, Gemini, MiniMax). MiniMax adds the sharper decomposition —
markets are (a) partially adversarial, (b) endogenous with respect to
the analyst's actions, and (c) non-stationary under regime shift, and
"each technique fails at a different rate" [T6] — uncited
in MiniMax (Source: MiniMax).
A second structural point, contested between reports and resolved in §6.9 below: MiniMax repeatedly discharges the endogeneity risk at this capital scale on the grounds that $100 sits "below the typical depth-1 visible quote on every CFTC-regulated economic/monetary contract." That universal quantifier is unevidenced — MiniMax's own digest flags it as having no order-book data behind it — and Gemini's venue audit contradicts it specifically, recording "thin depth on niche events" for Kalshi. The endogeneity discharge is retained only for high-volume contracts, not universally.
What the imports can and cannot do here. These
techniques are measurement and discipline instruments. Not one of them
generates edge. Proper scoring rules tell a forecaster whether stated
probabilities match realized frequencies; they do not make the
probabilities better. Alpha-spending controls the false-positive rate of
a self-assessment; it does not raise the win rate. Change-point
detection announces that an edge has decayed; it does not supply a
replacement. The honest characterization is that this entire toolkit
converts an unmeasurable 90-day outcome into a measured one, and
MiniMax, Gemini, and Qwen all separately conclude that the measured
answer over a 90-day, ≤30-trade sample will be indeterminate regardless
of technique quality. Gemini quantifies the ceiling: reaching t ≥ 3.0
over N = 90 trading days requires a daily Sharpe of 3/√90 = 0.3162, i.e.
an annualized Sharpe of 5.02 [T6] — Gemini's own
derivation, arithmetically verified in its digest but carrying no
citation — a figure with essentially no precedent in unleveraged retail
asset classes. Qwen states the same conclusion qualitatively: "Any live
result from this experiment would be considered statistical noise"
[T6] — uncited in Qwen.
(Source: MiniMax — 11 techniques; Gemini — 1 row; Qwen — 2 rows)
What the domain solves. Meteorology is the only
field that industrialized the separation of three properties finance
routinely collapses into one number [T6] — MiniMax's
framing, uncited (Source: MiniMax):
[T1].[T1].
(Gemini's JSON appendix renders this DOI as
10.1188/016214506000001437, contradicting its own markdown
— see Resolved Conflicts.)Transfer mechanism. A binary event contract is a
probability forecast with a cash settlement attached. MiniMax states
that transfer to event-contract pricing "is essentially free of
conceptual adaptation cost because event contracts are
probability estimates" [T6] — MiniMax's assertion, uncited.
The mechanism runs in three stages. First, hindcasting: score a trader's
hypothetical probabilities against historical settlements before capital
is risked, aggregating Brier and CRPS over a population of contracts
(MiniMax specifies a rolling window of N ≥ 50). Second, diagnosis: a
reliability diagram with bootstrap confidence bands (Bröcker & Smith
2007, Wea. & Fcst. 22(3): 651–661, doi:10.1175/WAF993.1
[T1]) determines whether the curve lies inside the 95% band
of the diagonal. Third, recalibration: Model Output Statistics (Glahn
& Lowry 1972, J. Appl. Meteor. 11(8): 1203–1211
[T1]) regresses raw model output onto historically observed
market mid-prices, which is a strict improvement whenever the raw output
is systematically biased.
Qwen proposes the narrowest and most immediately executable version
of the same mechanism: "tracking the Brier score of Kalshi's closing
prices for a set of resolved markets would provide a metric of how well
the market predicted those events" [T6] — uncited in Qwen.
That is a measurement of the venue, not of the trader, and it is the
cheapest diagnostic in this section — it requires only settled-contract
history and costs nothing.
Gemini adds the decomposition finance most often omits:
BS = REL − RES + UNC, the Murphy partition separating
reliability, resolution, and the irreducible uncertainty of the event,
which Gemini attributes to Brier (1950) and Gneiting & Raftery
(2007), doi:10.1198/016214506000001437 [T1]. MiniMax
supplies the continuous-score analogue,
CRPS = reliability + sharpness (Hersbach 2000, Wea.
& Fcst. 15(5): 559–570 [T1]), and notes CRPS "is
the only scoring rule that does not require binning choices that alter
the score" — relevant for multi-outcome events such as an FOMC rate-path
contract.
Transfer risk. The atmosphere is exogenous; the
quote is not. MiniMax frames it as endogeneity of the trade rule: any
calibration derived from historical venue ticks is contaminated by the
trader's own historical activity once size moves the book. Gemini states
it compactly: "Atmosphere non-adversarial; markets have adversarial
feedback." Qwen adds the generalization that a model performing well
in-sample on historical weather "may fail spectacularly if applied to
financial data, as the relationship it learned is not a causal law but a
transient statistical artifact" [T6] — uncited in Qwen.
Three additional domain-specific risks are stated: ensemble forecasting
assumes a physics-consistent multi-member ensemble that a retail
participant does not possess and must approximate by bootstrapping; MOS
recalibration warps over months and requires rolling
exponentially-weighted recomputation; reliability diagrams assume
forecasts are exchangeable across time and forecaster, which
autocorrelated forecasts from an evolving forecaster violate.
Gap. No source writes the actual formula for the Brier score, the log score, the Brier skill score, or the CRPS integral. MiniMax's digest flags this explicitly: the cluster designated as the authority on proper scoring rules never defines one. An implementer must source the definitions elsewhere.
Python. scores (Gemini, claimed v2.5.0,
version unverified [T6]) for CRPS/Brier/energy scores;
properscoring for brier_score,
crps_gaussian, crps_empirical,
threshold_brier_score — unmaintained, with the two reports
disagreeing on the last release date (Gemini: 2015-05-20; MiniMax: "0.1
released 2017"), both [T6]. MiniMax notes both Brier and
CRPS are "<30 lines each" implemented directly in NumPy, which
removes the dependency question entirely.
uncertainty-toolbox for reliability diagrams
[T6].
(Source: MiniMax — 11 techniques; Gemini — 1 row; Qwen — 5 named practices)
What the domain measured. The Good Judgment Project
ran inside IARPA-funded geopolitical forecasting tournaments and
produced four findings MiniMax describes as "not subject to dispute":
trained forecasters outperform aggregate analysts on Brier metrics at
1-week-to-1-year horizons; frequent updating is the dominant
behaviorally measurable contributor to score, not raw cognitive talent;
trimmed-mean and extremized aggregation beat simple averages under
proper scoring rules; and superforecaster performance stands in a
contested relationship to market-implied probabilities. MiniMax
attributes the first three to Mellers, Stone, Murray et al. (2015),
Persp. Psych. Sci. 10(3): 267–281, doi:10.1177/1745691615576804
[T1]; the fourth is contested and unresolved (see below).
Qwen names the same practice set from the other direction: frequent
belief updating, base-rate anchoring (outside-view reasoning),
decomposing questions into tractable components, thinking in shades of
gray, and combining diverse information sources [T2].
The replication record is the strongest claim of provenance in this
domain, and it is also where MiniMax overreaches. MiniMax asserts "at
least three independent meta-analyses from 2018–2023 reproduce the core
finding" but names only one — Himmelstein & Stahl (2023),
Judgment and Decision Making 18: e22, doi:10.1017/jdm.2023.23
[T2]. The other two are never identified. Treat the
replication claim as supported by one named systematic review, not three
[T6] on the "three meta-analyses" figure specifically.
Transfer mechanism. MiniMax supplies the only
explicit, fully parameterized trade rule anywhere in this section,
reported here descriptively as a source specification and not as a
recommendation: a daily-updated probabilistic ledger over N ≥ 50
candidate contracts, where "probability ≤ 0.10 ⇒ bet NO; 0.10 < p
< 0.90 ⇒ no trade; p ≥ 0.90 ⇒ bet YES, scaled as a fraction of stake
by min(p − ask, bid − p)." The aggregation mechanism is a linear opinion
pool over three inputs — prediction-market consensus, economist-survey
medians (e.g. the Survey of Professional Forecasters for macro), and
private signals — combined by trimmed mean (Clemen 1989, Int. J.
Forecasting 5(4): 559–583, doi:10.1016/0169-2070(89)90012-8
[T1]; Cooke 1981, Experts in Uncertainty
[T2]).
Two forms of extremizing appear, and they are not competing estimates of the same quantity. MiniMax specifies a linear blend shifting the private forecast toward the reference base rate by a fraction α ∈ [0.05, 0.15] per iteration. Gemini specifies logit extremizing with a scaling exponent d = 1.4 applied to an underconfident crowd or LLM consensus. These parameterize different operations on different scales; printing them adjacent invites a comparison that does not exist. Both are reported; neither adjudicates the other.
The frequency mechanism transfers most directly. MiniMax proposes
treating each macro data release as a "tournament tick" and updating
immediately rather than holding stale positions — the GJP finding that
update frequency, not talent, drives score. Qwen concedes that a retail
participant "cannot replicate the structured Delphi-style aggregation
used by the GJP" but argues the principles transfer: assign a prior,
update systematically, "rather than reacting emotionally to price
movements" [T6] — uncited in Qwen.
Base rate. MiniMax constructs the reference class
this problem actually needs — "will an unaffiliated retail trader with
documented edge turn $100 into $200 in 90 calendar days under
CFTC-regulated event contracts?" — from historical Polymarket
participants 2020–2025, Intrade retail proxy data 2004–2014, Betfair
economic-contract loss data, and the retail attrition literature, and
estimates it "on the order of 1–5%, not 10%." MiniMax labels this
[T6] — base-rate constructed by this author, and that label
is correct and is preserved. Qwen independently estimates "likely less
than 10% and perhaps much closer to 1%" with no citation whatsoever
[T6]. Two independent author constructions landing in the
same 1–5% region is weak convergent evidence, not a measurement; neither
report has individual-level participant P&L data, and MiniMax
explicitly names that as the verification that would be required.
Transfer risk. GJP experiments ran with (a)
low-stakes monetary incentives, (b) questions whose ground truth was
publicly observable in real time, and (c) participants pre-conditioned
by skill-selection panels. This problem inverts all three: the entire
$100 is at stake, the trader's own edge is the source of
probability information rather than a consumer of it, and no skill
selection has occurred. MiniMax names the dominant failure mode
calibration collapse under stake size — systematic
compression of probability estimates toward 0.5 when the bid-ask spread
is non-trivial, which is a direct violation of proper-scoring theory
[T6] — MiniMax's inference, uncited. Gemini names a
different and complementary risk: "GJP static long-horizon; order books
shift instantly on news" — the tournament questions resolved over weeks
to a year, whereas an event contract reprices in milliseconds. Qwen's
version is the reference-class problem: "Reference-class selection can
be arbitrary and introduce bias."
Contested finding — direction versus markets. Qwen
states superforecasters "consistently outperform both unstructured
groups and prediction markets." MiniMax states
superforecaster performance "converges with market-implied probabilities
on similar questions, but superforecasters move markets when they
update," and states the direction two different ways within its own
section. Gemini is silent. Neither directional claim carries adequate
sourcing — Qwen's rests on a Good Judgment Inc. self-published PDF (a
vendor-interested source), and MiniMax's rests on an Atanasov et al.
(2020) citation its own digest flags as having an implausible author
list, volume, and DOI for the title given. The directional claim is
excluded. What survives: structured aggregation of trained forecasters
outperforms unstructured individual judgment — Mellers et al. (2015),
doi:10.1177/1745691615576804, with review support in Himmelstein &
Stahl (2023), doi:10.1017/jdm.2023.23 [T1]. Whether it
beats a liquid market is unresolved in this evidence base.
Python. All scoring operations are implementable in
under ten lines of NumPy per MiniMax; properscoring or
scores cover them. scipy.optimize for logit
extremizing (Gemini). Elicitation and recalibration discipline is a
user-interface problem, not a library problem.
(Source: MiniMax — 11 techniques; Gemini — 1 row; Qwen — 3 techniques)
What the domain solves. Actuarial science optimized
a different objective than portfolio theory: not mean-variance of
returns but probability of ruin over a long operating horizon
with bounded premium income. MiniMax states the structural
correspondence precisely — "the doubling problem is literally
the dual of the actuarial problem: minimize P(ruin) on the path to a
target, given a fixed maximum loss budget, instead of minimize P(ruin)
on the path to insolvency, given a fixed maximum premium stream"
[T6] — MiniMax's framing, uncited. This is the deepest
conceptual import in the section, and it is the one all three reports
reach independently.
Transfer mechanism — ruin theory. Qwen supplies the
surplus process explicitly: U(t) = u + ct − S(t), with
u the initial surplus, c the premium income rate, and
S(t) the aggregate claims process; ruin probability
ψ(u) = P(inf_{t ≥ 0} U(t) < 0) — the classical
Cramér–Lundberg formulation, given a modern treatment in Asmussen &
Albrecher (2010), Ruin Probabilities [T2]. Gemini
supplies the bound: Lundberg's inequality,
ψ(u) ≤ e^(−Ru), where the adjustment coefficient R
is the unique positive root of λ + cR = λM_X(R) — same
lineage; Asmussen & Albrecher (2010) [T2].
Operationally, the trader's premium income is the realized mean of the
strategy's log-return distribution per normalized period and the claim
size is the realized loss distribution; the Lundberg coefficient
resolves whether a positive expected log-return is sufficient to make
P(ruin) < 1.
MiniMax prints a version of the adjustment-coefficient condition — "E[e^(γX)] < 1 for some γ > 0" — that is unsatisfiable for any positive claim size, since e^(γX) > 1 pointwise for X > 0 and every γ > 0. Its own digest flags this (CI-2) and further notes it is falsely attributed to Asmussen & Albrecher. The malformed condition is excluded; Gemini's and Qwen's correct forms are carried.
Transfer mechanism — frequency-severity
decomposition. All three reports name this as the correct first
decomposition of any candidate return-generating process (Panjer 1981,
ASTIN Bulletin 12(1): 22–26, doi:10.1017/S0515036100006615
[T1]). MiniMax works the example for a
long-out-of-the-money weekly index call: severity is the right tail of
the log-return distribution, where kurtosis dominates and the log-normal
right tail is far too thin; frequency is the count of independent
observations per quarter under a purged, cross-validated
effective-sample count. Panjer recursion then computes the exact
aggregate distribution S = X₁ + … + X_N for small trade
counts — MiniMax specifies K = 10 trades per quarter — giving a
finite-horizon distribution of compounded P&L rather than an
asymptotic approximation. That exactness matters here precisely because
the sample is tiny; asymptotic approximations are worthless at N = 10.
MiniMax's kurtosis figure ("~10–20 for daily log-returns of liquid US
equities") is stated with no source, no sample period, and no universe
definition [T6].
Transfer mechanism — credibility theory. Bühlmann
(1967), ASTIN Bulletin 4(3): 199–207,
doi:10.1017/S0515036100008832 [T1], and its heteroskedastic
extension Bühlmann & Straub (1970), Mitt. Ver. Schweiz.
Versicherungsmathematiker 70: 111–133 [T1], answer the
question a backtest cannot: what is the prior probability that a
strategy has real edge, given that it appears in the literature at all?
The credibility factor Z = n/(n + K) shrinks a
strategy-edge estimate toward the population mean of pre-registered
retail strategies, with weight rising in observation count n.
All three reports name this technique; Gemini adds the operational
framing that it "blends backtest alpha with retail base rates," Qwen
that it "blends historical data with prior expectations to form more
reliable estimates." At N ≈ 10–30 trades, Z is small and the shrinkage
is severe — which is the correct behavior and also the reason a 90-day
live result cannot escape its prior.
Extreme-value theory. MiniMax specifies fitting the
empirical peaks-over-threshold Generalized Pareto Distribution on each
candidate strategy's worst 5% of observations and verifying GPD fit
before trusting any estimated Sharpe (Embrechts, Klüppelberg &
Mikosch 1997 [T2]; McNeil, Frey & Embrechts 2015
[T2]). This technique appears in MiniMax's body but not in
its Table D — one of roughly twelve such omissions its digest flags.
Explicitly non-transferable. MiniMax cites Mack (1993) chain-ladder loss-development triangles and then disclaims it: "I cite only because the reader is likely to encounter it. It's not directly applicable here." Its consolidated risk table marks it "Largely irrelevant; do not transfer." That disclaimer is preserved as stated — it is the correct behavior for a technique inventory and the only instance of it in any of the three reports.
Transfer risk. Ruin theory assumes claim sizes are i.i.d. and exogenous to the insurer's activity. Trading returns are serially correlated through overnight gaps and macro cycles, and the distribution is non-stationary. MiniMax names the subtler failure: ruin theory does not condition on the data-generating process changing in response to the analyst's signal, so where the signal correlates with the future evolution of the distribution, ruin estimates are systematically optimistic. Gemini's compact form: "Claims assume i.i.d. independence; asset returns exhibit tail clustering." Qwen: credibility theory "assumes the underlying risk process is stationary, often false in financial markets." Bühlmann credibility additionally assumes mutually independent risk classes, whereas candidate strategies are correlated through shared macro factors; the stated mitigation is multi-level credibility or random effects. MiniMax's discharge of the endogeneity risk at $100 scale ("the transfer risk for actuarial ruin theory at this scale is therefore manageable") rests on the same unevidenced depth claim addressed in §6.9.
Python. lifelib (actuarial projection
primitives, version unverified [T6]);
chainladder (loss development, [T6], and
relevant only to the technique MiniMax disclaims); Lundberg exponent by
direct numerical solution in ~50 lines of
numpy/scipy.optimize; Panjer recursion in ~30
lines; Monte Carlo ruin estimation in numpy.random.
(Source: MiniMax — 8 techniques; Gemini — 1 row; Qwen — no coverage)
What the domain solves. Nowcasting estimates the
current value of a latent quantity from sparse, delayed, and incomplete
observations. MiniMax states three transferable findings: hierarchical
Bayesian pooling across subpopulations produces a latent-state posterior
that tracks truth far faster than raw data does; reporting-delay
distributions combined with current reported counts yield the nowcast
without individual-level record linkage; and backfill
correction — re-estimating past incidence as later reports
complete — is empirically necessary, because uncorrected past estimates
are systematically low — Höhle & an der Heiden (2014),
Biometrics 70(4): 993–1002, doi:10.1111/biom.12194, generalized
in Günther et al. (2021), doi:10.1002/bimj.202000112 [T1].
MiniMax's "tracks truth within hours while raw data lags by weeks" is
stated as a general property with no study, disease, or metric attached
[T6].
Transfer mechanism. The macro-release calendar is
the direct analogue. A trader faces CPI, NFP, PCE, and FOMC releases at
irregular intervals with information leaking between them through Fed
speeches, equity returns, and survey data. The mixed-frequency framework
of Giannone, Reichlin & Small (2008), J. Monetary Econ.
55(4): 665–676, doi:10.1016/j.jmoneco.2008.05.010 [T1] —
itself the paper that imported "nowcasting" into macroeconomics —
produces a daily-updated estimate of the unobserved macro state.
MiniMax gives the only fully specified state-space model in the section: state vector = (latent inflation nowcast, latent unemployment nowcast, latent recession probability); observation vector = (released CPI, released NFP, market-implied probabilities from Kalshi/CME FedWatch); transition dynamics = AR(1) latent drift. The posterior mean is a daily probability surface over questions such as "will CPI exceed 3.0% YoY at the next release?", which is then priced against the corresponding event contract. This is the most directly executable specification in Section 6 and requires no proprietary data.
Gemini contributes the operational package and a distinct originating
citation: NobBS Bayesian delay nowcasting, McGough et al.
(2020), PLOS Computational Biology 16(4): e1007735,
doi:10.1371/journal.pcbi.1007735 [T1], applied to backfill
correction on BLS/GDP/CPI releases. MiniMax cites the methodological
origin instead — Höhle & an der Heiden (2014), Biometrics
70(4): 993–1002, doi:10.1111/biom.12194 [T1], generalized
by Günther et al. (2021), Biometrical Journal 63(8): 1575–1593,
doi:10.1002/bimj.202000112 [T1]. These are complementary,
not competing: Höhle & an der Heiden is the originating Bayesian
nowcasting method, McGough et al. the widely used implementation. Both
are carried.
Hierarchical pooling across venues. MiniMax proposes
a partial-pooling model with venue-specific intercepts and a common
latent-state loading across Kalshi, ForecastEx, and IBKR event
contracts, which yields a strictly better probability estimate than any
single venue when the venues are partially segmented —
"operationally the same as the CDC's pooling of test-positivity across
states" [T6] — MiniMax's analogy, uncited. The conditional
matters: MiniMax's own risk table notes that cross-market arbitrage
collapses the mispricings the pooling is meant to exploit, so the
technique's value is inversely proportional to how integrated the venues
are.
Backfill correction applied to the ledger. Each closed position's realized outcome is re-fed into a hierarchical model of the trader's own calibration, so that a Monday probability revealed as wrong by Tuesday's release is retro-corrected before it enters the calibration record. This is the epidemiological insight that most directly attacks the small-sample problem in §6.0: it extracts more information per settled contract than naive scoring does.
Transfer risk. Epidemiological nowcasting assumes (a) the reporting system is exogenous, (b) the data-generating process is approximately stationary, and (c) the analyst cannot influence the data. MiniMax judges the first two to transfer in a bounded way during normal regimes and the third not to transfer cleanly at material size. Gemini names a risk MiniMax does not, and it is the sharper one: "Clinical delays are physical; economic data strategically revised." Reporting delay in an outbreak is a physical and administrative lag. Macro revision is a decision made by an agency with its own objectives and calendar, and it can move in either direction. A second failure mode: nowcasting posterior variance depends on correct noise-model specification, and macroeconomic noise is a sum of measurement error, seasonal effects, and revisions — omitting any one produces overconfident intervals. The stated remedy is mandatory reporting-delay modeling as a hierarchical prior on the error variance.
Coverage note. Qwen omits this domain entirely.
Python. PyMC (Gemini claims v5.17.0,
unverified [T6]) with an explicit reporting-delay layer;
arviz for posterior diagnostics;
cmdstanpy/Stan for HMC-NUTS (Carpenter et al. 2017, J.
Stat. Software 76(1), doi:10.18637/jss.v076.i01 [T1]);
filterpy.kalman and statsmodels.tsa.statespace
for the Kalman specification; MiniMax notes the Kalman filter is ~100
lines of NumPy written directly.
(Source: MiniMax — 12 techniques; Gemini — 1 row; Qwen — 1 technique, SPRT)
What the domain solves. MiniMax maps three decision problems: strategy-level regime detection (when has the edge regressed to zero?), position-level online learning (posterior belief about a latent variable), and sequential decision-making under explicit type-I/type-II error budgets. This is the domain with the most techniques and the most immediate operational bite, because it addresses the question a 90-day experiment must answer continuously — is this still working?
Change-point detection. CUSUM (Page 1954,
Biometrika 41(1/2): 100–115 [T1]) runs on the
running expected log-return, or equivalently on running P&L
normalized by per-trade risk; once the cumulative sum exceeds a
threshold tuned via in-control Average Run Length, the strategy is
declared drifting. MiniMax makes the asymmetry argument that matters: in
a non-stationary environment CUSUM is conservative — false alarms too
rare — which is the correct direction of error for capital protection.
The GLR variant (Lorden 1971, Ann. Math. Stat. 42(6):
1897–1908, doi:10.1214/aoms/1177693014 [T1]) estimates the
post-change parameter rather than committing to a fixed target, and
MiniMax recommends it specifically because the post-degradation
parameter is unknown ex ante.
Bayesian Online Change-Point Detection (Adams & MacKay 2007,
arXiv:0710.3742 [T3]; refereed treatment Fearnhead &
Liu 2007, JRSS B 69(4): 589–605,
doi:10.1111/j.1467-9868.2007.00545.x [T1]) returns a
posterior over run length, updating in O(N) per step and behaving
acceptably at small sample sizes. MiniMax's operationalization: feed
daily P&L in with a hazard rate tuned to expected strategy
half-life; the posterior P(run length > k) is the strategy's
instantaneous credibility. Gemini applies the same technique one level
down, to order-book regime shifts and volatility breaks for stop-out
triggering.
Sequential testing. Wald's SPRT (1945, Ann.
Math. Stat. 16(2): 117–186, doi:10.1214/aoms/1177731118
[T1]) is the single technique all three reports name.
MiniMax parameterizes it: test H₀ (win rate = 50%) against H₁ (win rate
= 60%) at α = 0.05, β = 0.20; for a true 60% win rate the test
terminates on average after ~30 trades, while under the null it nearly
always runs to its upper bound. The ~30-trade figure is stated without
formula, parameters, or derivation [T6], though the Average
Sample Number is computable in principle. Wald & Wolfowitz (1948),
Ann. Math. Stat. 19: 326–329, doi:10.1214/aoms/1177699121
[T1] proved SPRT minimizes expected sample size among all
tests with the same error rates — the result that also underwrites group
sequential clinical-trial design in §6.8.
The 30-trade figure deserves emphasis against §6.0: the number of trades required to distinguish a 60% win rate from a coin flip is roughly the same order as the total number of trades a 90-day, $100 experiment can execute under T+1 settlement. The experiment is at the resolution boundary of its own test.
Filtering. The Kalman filter (Kalman 1960, J.
Basic Eng. 82(1): 35–45, doi:10.1115/1.3662552 [T1])
estimates log-volatility as a latent state from windowed returns;
MiniMax argues the smoothed mean beats rolling standard deviation
because it adapts the smoothing constant to the noise-to-variance ratio.
The particle filter (Gordon, Salmond & Smith 1993, IEE Proc.
F 140(2): 107–113, doi:10.1049/ip-f-2.1993.0014 [T1])
handles non-Gaussian states such as discrete regime membership — ~150
lines for a one-dimensional state. Extended and unscented variants
(Julier & Uhlmann 1997, 2004, Proc. IEEE 92(3): 401–422,
doi:10.1109/JPROC.2004.823170 [T1]) appear in MiniMax's
body without a stated transfer risk.
Control charts. Shewhart (1924; Montgomery 2019
[T2]) and EWMA (Roberts 1959, Technometrics 1(3):
239–250, doi:10.1080/00401706.1959.10489860 [T1]) are the
cheapest instruments here — EWMA is a single smoothed
deviation-from-target with two-sigma bands, and a breach declares regime
change. MiniMax's own risk table notes Shewhart has low power against
small shifts and EWMA is sensitive to the smoothing-parameter choice, so
the low cost buys correspondingly low resolution.
Transfer risk. MiniMax names three distinct failures, and this is the most rigorous transfer-risk treatment in any domain:
Gemini names a fourth: "Signal processing assumes Gaussian white noise; returns feature jump diffusion" — which is precisely why MiniMax's escalation from Kalman to particle filtering is the correct response rather than an optional refinement.
Python. ruptures for change-point
detection (Gemini claims v1.1.9, MiniMax "1.x range," both
[T6]; MiniMax's cited API path
ruptures.detect.cusum does not exist — the package exposes
search classes Pelt, Binseg,
Window, BottomUp, Dynp with cost
functions); filterpy for Kalman and particle filters;
statsmodels.stats.diagnostic.breaks_cusumolsresid;
arch for volatility models;
bayesian-changepoint-detection for BOCPD (existence
unverified [T6]).
(Source: MiniMax — 9 techniques; Gemini — 1 row; Qwen — no coverage)
What the domain solves. Item-response theory
estimates latent ability where items are noisy, ability is only
partially observable, and the object of inference is the
responder, not the population. MiniMax states the mapping: this is
the same problem as crediting a prediction source — a forecaster, an
indicator, an NLP sentiment model, an FOMC statement — with empirical
reliability [T6] — MiniMax's mapping, uncited.
Transfer mechanism — sources as raters. Each
candidate signal is treated as a rater and each contract resolution as
an item. Dawid & Skene (1979), JRSS C 28(1): 20–28,
doi:10.2307/2346806 [T1] estimate both latent truth and
per-rater error rates by EM. MiniMax's operationalization: treat each
"signal predicts YES on contract X" as a binary rating against ground
truth, estimate rater accuracies over a sliding window of N = 100
contracts, and use the latent-truth estimate as the pooled probability.
The hierarchical rater model (Patz, Junker, Johnson & Mariano 2002,
ETS Research Report [T5]) extends this by clustering
signals by source with source-level and contract-level parameters.
Transfer mechanism — the trader as examinee. Running
Rasch (1960 [T2]) or 2PL IRT (Birnbaum 1968, in Lord &
Novick [T2]) on the trader's own forecast ledger treats
each historical forecast as an item and the trader's calibration as a
single latent-trait parameter, jointly estimating item difficulty and
discrimination. The output is a calibration estimate that properly
accounts for the difficulty of the questions faced — which naive
Brier scoring does not. Gemini specifies the 3PL variant, adding a
guessing parameter c_j alongside skill θ and difficulty b_j (Rasch 1960
/ Lord 1980 [T2]/[T3]), which is the correct
structure for binary contracts where a coin flip scores 50%.
Unification claim. MiniMax asserts that "Bühlmann
credibility is a special case of IRT with a Rasch model whose item
parameters are pooled," unifying §6.3 and §6.6 [T6] —
stated without proof or citation. If correct, the actuarial shrinkage
and the psychometric ability estimate are the same estimator viewed from
two domains, and an implementer needs only one of them. The claim is
stated without proof or citation and should be treated as the source's
synthesis rather than an established result.
Generalizability theory (Cronbach, Gleser, Nanda
& Rajaratnam 1972 [T2]) decomposes reliability into
within-source, between-source, and item-heterogeneity components —
separating "this strategy is fragile to the choice of source" from "this
strategy's signal quality is genuinely high." MiniMax covers it in the
body without a stated transfer risk.
Transfer risk. Three named by MiniMax, none carrying
a citation [T6]:
Gemini adds a fourth: "Psychometric traits stable; trader skill fluctuates." A student's latent ability is approximately constant across a test session. Trading skill is state-dependent — on stake size, on fatigue, on regime — which violates the core exchangeability assumption more severely than concept drift in a source does.
Coverage note. Qwen omits this domain entirely.
Python. pyirt (version unverified
[T6]); Dawid–Skene EM in ~30 lines of direct implementation
per MiniMax; PyMC for the hierarchical rater model;
scipy.optimize for a custom 3PL (Gemini);
factor_analyzer as an adjacent tool.
statsmodels does not ship IRT.
(Source: MiniMax — 8 techniques; Gemini — 1 row; Qwen — no coverage)
What the domain solves. Entropy is the canonical
measure of uncertainty and mutual information the canonical measure of
association. The result that matters here is Kelly (1956), Bell
System Technical Journal 35(4): 917–926,
doi:10.1002/j.1538-7305.1956.tb03809.x [T1], which
identifies the maximum achievable exponential growth
rate of capital with the mutual information between the
bettor's private signal and the realized outcome.
A formula excluded. MiniMax states, attributing it
to Cover & Thomas (2006), that "the growth-rate-optimal bet fraction
is f* = I(X; Y) / H(X)." This is a category error and
MiniMax's own digest flags it as such (CI-1): Kelly's identity equates
the achievable growth rate — a quantity in bits or nats per bet
— with the mutual information. The optimal bet fraction is a
dimensionless capital share and a different object entirely. The formula
is repeated in MiniMax as the basis for its signal-selection
prescription. It does not enter this report. The
correct statement — growth rate equals mutual information under optimal
play — is carried; the sizing rule derived from it is not.
Transfer mechanism — signal selection. Stripped of
the bad formula, the usable prescription survives: for two candidate
signals with mutual information I(Y; S₁) and I(Y; S₂) about the outcome
under the same capital budget, the higher-mutual-information signal
supports a higher growth rate. MiniMax argues this is strictly better
than ranking by raw predictive accuracy, because a low-accuracy but
high-conditional-MI signal can outperform a high-accuracy but
near-redundant one — the redundancy is invisible to accuracy and visible
to MI. MiniMax marks the finance-specific literature for this criterion
[T6] ("dispersed across practitioner conference
proceedings") and points back to Kelly (1956) as the canonical anchor,
which is the correct handling.
Transfer mechanism — divergence as contract screen.
For a binary contract with subjective probability p̂ and market mid-price
m, the Kullback–Leibler divergence D_KL(p̂ ‖ m) (Kullback
& Leibler 1951, Ann. Math. Stat. 22: 79–86,
doi:10.1214/aoms/1177729694 [T1]) relates to expected
log-growth over the contract. Gemini names the same object alongside
channel capacity I(X;Y).
This screen carries a trap that must be stated with it. MiniMax writes that selecting contracts where p̂ and m diverge significantly "is mathematically equivalent to maximizing log-growth per dollar." The identity holds only when p̂ is the true probability. Where p̂ merely differs from m, large divergence signals large expected loss exactly as readily as large expected gain. As written, the sentence licenses "bet wherever you disagree with the market" — which is the precise failure the §6.1 calibration machinery exists to prevent. MiniMax's own digest flags this (CI-16). The screen is retained strictly as a second filter downstream of demonstrated calibration, never as a standalone entry criterion.
Entropy pooling (Meucci 2010 [T4])
projects a subjective view onto a market-implied prior by minimum
relative entropy — a generalization of Black-Litterman admitting
arbitrary constraints. Maximum-entropy priors (Jaynes
1957, Physical Review 106: 620–630, doi:10.1103/PhysRev.106.620
[T1]) supply a least-committal distribution consistent with
known moments, useful where a signal source is genuinely model-free.
Fano's inequality bounds misclassification probability
given mutual information, yielding a lower bound on the risk of the
trader's bets. MiniMax cites Fano (1961), Transmission of
Information, MIT Press, in its Table D and body but records no
bibliographic entry for it anywhere [T6] on the citation
trail.
Transfer risk. Three named by MiniMax; the first two
carry no citation [T6], the third names the KSG and
Miller–Madow corrections for which MiniMax records no bibliographic
entry [T6] on the citation trail:
[T2]) is
markedly more delicate and MiniMax states it is rarely advisable without
simulation.[T6], and if correct it means a naive MI screen at this
sample size will nominate signals carrying no information. MiniMax names
both KSG and Miller–Madow with no bibliographic record for either
[T6] on the citation trail.There is a structural tension between this domain and the problem, noted by MiniMax and by Qwen's Cluster 1 material: Kelly maximizes long-run geometric growth over many periods, which is not the objective of a fixed-multiple target under a hard deadline. Both reports treat Kelly as the indispensable starting reference and the wrong optimand for this specific problem.
Coverage note. Qwen omits this domain entirely.
Python. dit for discrete information
theory; sklearn.feature_selection.mutual_info_classif
(biased — MiniMax explicitly directs to a KSG implementation for serious
use); scipy.stats.entropy (Gemini);
cvxpy/scipy.optimize for entropy pooling (~50
lines per MiniMax); numpy for the KL screen.
(Source: MiniMax — 11 techniques; Gemini — 1 row; Qwen — no coverage as a distinct domain; Qwen assigns SPRT to industrial statistics)
What the domain solves. MiniMax maps three decision problems with unusual precision: when is enough evidence accumulated to declare a treatment effective (when is the edge real?); how do you monitor for harm without prematurely abandoning a useful treatment (when is failure rate elevated enough to stop?); and how do you control the false-positive rate when peeking at the data every day (how often can a trader check P&L and still trust the verdict?). The third is the one no other domain in this section addresses, and it is the one a 90-day experiment with continuously visible P&L needs most.
Transfer mechanism — alpha spending. Under naive
repeated testing at α = 0.05 across K looks, the cumulative type-I error
inflates. MiniMax specifies K = 13 weekly checks over 90 days and
computes 1 − (1 − 0.05)¹³ ≈ 0.49. The arithmetic is right (0.95¹³ ≈
0.513) but the bound is for 13 independent tests;
interim looks at accumulating data are strongly positively correlated,
so true inflation from 13 sequential looks is materially lower than
0.49. MiniMax's own digest flags this (CI-3). The direction of the
argument survives — repeated unstructured peeking at P&L inflates
false-positive rates substantially — but the 0.49 figure overstates the
magnitude and is marked [T6].
The corrective machinery is well established and both MiniMax and
Gemini cite it identically, which is the cleanest cross-report agreement
in Section 6: Pocock (1977), Biometrika 64(2): 191–199,
doi:10.1093/biomet/64.2.191 [T1] for group sequential
design with equal α per look; O'Brien & Fleming (1979),
Biometrics 35(3): 549–556, doi:10.2307/2530245
[T1] for the conservative-early boundary; and Lan &
DeMets (1983), Biometrika 70(3): 659–663,
doi:10.1093/biomet/70.3.659 [T1] for the continuous
alpha-spending function that removes the requirement to fix the number
of looks in advance. Gemini's framing: "Controls cumulative FDR at α =
0.05 across interim reviews."
Calendar time versus information time. MiniMax makes
an argument no other report makes and it is the most operationally
consequential item in this subsection. With 90 days and probably ≤30
trades, calendar-time fraction (days elapsed / 90) and information-time
fraction (effective sample size / target ESS) diverge sharply, because
serial correlation makes the effective sample smaller than the trade
count implies. Lan & DeMets (1989), Stat. in Medicine
8(10): 1191–1198, doi:10.1002/sim.4780081003 [T1] show
information-time spending is preferable for heterogeneous designs.
MiniMax concludes this problem should use information-time
alpha-spending. The practical effect: a trader 60 days into the
experiment has spent far less than two-thirds of the available alpha,
because the information accumulated is less than the calendar
suggests.
Haybittle–Peto. Only the final look is evaluated at
nominal α; earlier interim looks use α_interim ≈ 0.001. MiniMax argues
this is appropriate here because the cost of an early false positive —
declaring the strategy works and increasing risk on the strength of
noise — dominates the cost of a delayed decision. MiniMax marks its own
appropriateness claim [T6] and notes the Haybittle (1971)
and Peto et al. (1976) citations require confirmation; neither has a
complete bibliographic record in the file.
Pre-registration. MiniMax specifies three mandated
components: the precise hypothesis with all parameters bound (its worked
example: "long-volatility-on-CPI-day strategy produces positive expected
log-return under a 60-day horizon with N = 20 trades, loss-cap of $50,
expected win rate ≥ 55%"), the stopping rule as an a priori
alpha-spending function, and the look-elsewhere correction recording how
many candidate strategies were considered before this one. That third
component is the one retail practice universally omits and the one that
determines whether the final result means anything. CONSORT (Schulz,
Altman & Moher 2010, BMJ 340: c332, doi:10.1136/bmj.c332
[T1]) supplies the reporting standard.
Bayesian sequential design. Declare success when the
posterior P(edge > 0 | data) > 0.95 under a beta-binomial
conjugate setup (Spiegelhalter, Abrams & Myles 2004
[T2]). MiniMax calls this "the simpler and operationally
cleaner alternative to alpha-spending for a single trader," and on a
technical assessment it is: the conjugate posterior is a two-line
computation, requires no boundary tables, and handles unscheduled looks
natively.
Transfer risk. Four named:
[T6] — MiniMax's judgment, uncited. Institutional
pre-registration cost is amortized across millions of dollars; at $100
the human-attention cost is disproportionate. MiniMax's resolution is
precise: "The transfer is at the discipline, not at the
registry." A written protocol before launch, not a registry
submission.The meta-risk this domain names and no other does. MiniMax's §23.4: high methodological sophistication — proper scoring, credibility weighting, alpha spending — creates the illusion of robust edge, because every score is good, every credibility factor is high, and every interim peek passes. This is a documented failure mode in clinical trials, the demonstration of benefit rather than of true effect. The remedy is identical in both domains: at conclusion, perform one final fully-specified test, and accept that the answer can be "no edge" without the prior work having been wasted. A trader who has implemented every technique in Section 6 and reaches day 90 with $61 has run a successful experiment and an unsuccessful strategy, and the methodology's job is to make those two statements distinguishable.
Coverage note. Qwen omits this domain as a distinct source; its SPRT row is attributed to industrial statistics.
Python. gsDesign and rpact
(both R; MiniMax recommends porting boundary computation to Python);
scipy.stats.norm for spending functions (Gemini);
scipy.stats.beta for exact beta-binomial conjugate
posteriors; PyMC for the fuller Bayesian design. MiniMax's
§22.5 names betaind "from scipy" — no such symbol
exists in SciPy; its own Table D correctly says
scipy.stats.beta.
The single structural source. Financial markets are
(a) partially adversarial, (b) endogenous with respect to the analyst's
own actions, and (c) non-stationary under regime shift. Every technique
in this section assumes at least one of: an exogenous data-generating
process, analyst actions that do not move the process, or an
approximately stable distribution. Each technique fails at a different
rate against these violations [T6] — convergent across all
three reports, primary citation supplied by none (Source: MiniMax,
Gemini, Qwen).
The resolved conflict on scale. MiniMax discharges the endogeneity risk across four separate domains (§15.4, §17.4, §18.4, §23.3) on the grounds that a $100 position sits below depth-1 on "every" CFTC-regulated economic/monetary contract, concluding the transfer risk is "negligible at this scale and dominant at institutional scale." Its own digest flags the universal quantifier as unevidenced — no order-book data supports it — and one instance of the sentence is printed in self-negating form ("negligible at scale and dominant at scale"). Gemini's venue audit contradicts it for the specific case that matters: Kalshi is recorded as having "thin depth on niche events," and Gemini's friction matrix assigns Kalshi a total drag of 3.00%–15.00%, with the ceiling function producing "10% drag on a single $0.10 bet." The resolution carried here: endogeneity risk is plausibly small on high-volume contracts and is not established on thin ones, and thin niche contracts are precisely where a mispricing screen of the kind §6.7 describes will direct attention. The discharge does not generalize.
Ranking the failure modes by how much they bind at N ≈ 10–30 trades. Three independent domains produce the same wall from different directions, and the convergence is the most decision-relevant finding in this section:
| Constraint | Domain | Binding form at this scale |
|---|---|---|
SPRT terminates on average after ~30 trades for a true 60% win rate
[T6] on the figure |
Signal processing (§6.5) | The test needs roughly the whole experiment to resolve |
| Credibility factor Z = n/(n+K) is small at n ≈ 10–30 | Actuarial (§6.3) | The posterior barely moves off the prior |
| IRT/EM is poorly identified below N ≈ 30 items | Psychometrics (§6.6) | Ability estimate is unstable without strong priors |
| MI estimated from small samples is biased upward | Information theory (§6.7) | Signal-value claims are inflated exactly when least verifiable |
t ≥ 3.0 over 90 days requires annualized Sharpe ≥ 5.02
[T6] |
Backtesting integrity (Gemini) | No unleveraged retail asset class supports this |
These are not five risks. They are one risk — the sample is too small to support inference — measured by five instruments that were designed in five different fields and that agree. The asymmetry is important and runs one direction: this sample size can reject a strategy (a large loss is informative) far more readily than it can confirm one (a doubled stake is not). Any technique in this section that returns a favorable verdict at N ≈ 20 should be read as uninformative rather than as supportive.
The meta-risk. Across the imported inventory, the most underappreciated risk is the existence of well-capitalized counterparties whose actions appear in the price. No other domain in this section has this property: no weather system responds to a forecaster's prediction, no forecasting-tournament question moves because a participant updated, case counts are physical, and a manufacturing line has no opinion about the control chart. (Source: MiniMax)
The meta-meta-risk. Methodological sophistication manufactures false confidence. A full implementation of Section 6 produces a dashboard on which every number looks good — and none of those numbers is a measurement of edge. Only the final, pre-specified test is. (Source: MiniMax)
Domain labels are normalized to the eight named source domains; where
a source used a different label (MiniMax: "Macro-econometrics,"
"Industrial statistics," "Information theory / gambling"), the
normalized label is used and the source's label noted. Techniques
appearing in more than one source domain are merged into a single row
with the dual attribution stated. GAP marks a technique for
which the naming source stated no transfer risk; those cells are not
filled in. Python versions are unverified [T6] throughout —
no source's version claims were independently confirmed, and both
digests flag their package metadata as unauditable.
| # | Technique | Source domain | Originating citation | Transfer mechanism to this problem | Transfer risk | Python implementation |
|---|---|---|---|---|---|---|
| 1 | Brier score | Meteorology | Brier (1950), Mon. Wea. Rev. 78(1): 1–3,
doi:10.1175/1520-0493(1950)078<0001:VOFEIT>2.0.CO;2
[T1] (Qwen, Gemini, MiniMax) |
Score subjective probabilities against binary event-contract settlements; aggregate over rolling window of N ≥ 50 contracts before risking capital | Assumes forecaster cannot move the outcome; market prices treated as rational, so systematic biases are absorbed as noise. Endogenous once size moves the book | scores, properscoring.brier_score, or ~5
lines NumPy |
| 2 | Brier skill score (BSS) | Meteorology | Brier (1950); formalized in Murphy (1973), Mon. Wea. Rev.
101(7): 603–608,
doi:10.1175/1520-0493(1973)101<0603:HATMOT>2.0.CO;2
[T1] (MiniMax) |
Normalize Brier against a reference climatology to express skill relative to a naive forecast | GAP — no BSS-specific transfer risk stated; source
covers it only under the shared proper-scoring-rule row |
Manual; Brier plus reference baseline |
| 3 | Murphy score decomposition, BS = REL − RES + UNC |
Meteorology | Brier (1950) / Gneiting & Raftery (2007), JASA
102(477): 359–378, doi:10.1198/016214506000001437 [T1]
(Gemini) |
Separate reliability, resolution, and irreducible event uncertainty in event-contract pricing, isolating which component the trader actually controls | Atmosphere non-adversarial; markets have adversarial feedback loops | scores; manual partition |
| 4 | Logarithmic score (log-loss) | Meteorology / information theory | Good (1952), JRSS B 14(1): 107–114,
doi:10.1111/j.2517-6161.1952.tb00085.x [T1] (MiniMax;
named uncited by Qwen) |
Score subjective binary probabilities with a proper rule that penalizes hedging harder than Brier | Near-infinite penalty when p̂ → 0 and the outcome occurs; one mis-stated near-certainty dominates the record | sklearn.metrics.log_loss; NumPy |
| 5 | Continuous Ranked Probability Score (CRPS) | Meteorology | Matheson & Winkler (1976), Management Science 22(10):
1087–1096, doi:10.1287/mnsc.22.10.1087; decomposition in Hersbach
(2000), Wea. & Fcst. 15(5): 559–570 [T1]
(MiniMax) |
Score distributional forecasts for multi-outcome events (e.g. posterior over the FOMC rate path); no binning choice can alter the score | Fitted predictive distribution may not match the actual; reference climatology non-stationary | properscoring.crps_empirical,
crps_gaussian; scores |
| 6 | Sharpness–calibration decomposition | Meteorology | Gneiting, Balabdaoui & Raftery (2007), JRSS B 69(2):
243–268, doi:10.1111/j.1467-9868.2007.00543.x [T1]
(MiniMax) |
Decompose CRPS into reliability + sharpness to diagnose a forecaster who is calibrated but uninformative | True distribution is unobservable and non-stationary, so the decomposition's reference drifts | Manual computation on CRPS sub-components |
| 7 | Proper scoring rule convention | Meteorology | Gneiting & Raftery (2007), JASA 102(477): 359–378,
doi:10.1198/016214506000001437 [T1] (MiniMax,
Gemini) |
Constrain in-strategy loss functions to the class that rewards honest probability reporting | No external scorer enforces propriety on a solo trader; discipline is internal only | Implement scoring rules directly |
| 8 | Reliability diagram with bootstrap confidence bands | Meteorology | Bröcker & Smith (2007), Wea. & Fcst. 22(3):
651–661, doi:10.1175/WAF993.1 [T1] (MiniMax) |
Visual test of whether the calibration curve lies inside the 95% band of the diagonal before further capital deployment | Assumes forecasts exchangeable over time and forecaster; real forecasts are autocorrelated and the forecaster evolves. In-sample fit risk | uncertainty-toolbox; matplotlib + bootstrap bands |
| 9 | PIT / rank histogram | Meteorology | Dawid (1984), Z. Wahrsch. verw. Gebiete 60: 305–313,
doi:10.1007/BF00524500; Hamill (2001), Mon. Wea. Rev. 129(3):
550–560 [T1] (MiniMax) |
Diagnose distributional calibration for non-binary events (rate paths, index levels) | Same exchangeability violation as row 8 | Manual histogram + KS test |
| 10 | Ensemble forecasting | Meteorology | Leith (1974), Mon. Wea. Rev. 102(6): 409–418,
doi:10.1175/1520-0493(1974)102<0409:TSOMCF>2.0.CO;2
[T1] (MiniMax); Qwen attributes to "Wilks (2011)"
with a falsified URL |
Bootstrap N draws of the signal distribution to quantify output uncertainty and tail risk; e.g. 1,000 simulated paths for a CPI release | Assumes a physics-consistent multi-member ensemble that a retail participant does not have; bootstrap sampling distribution may not match true uncertainty. Financial feedback loops correlate errors across members | numpy.random bootstrap; conformal bands as sanity
check |
| 11 | Model Output Statistics (MOS) | Meteorology | Glahn & Lowry (1972), J. Appl. Meteor. 11(8): 1203–1211
[T1] (MiniMax) |
Regress raw model output onto historically observed venue mid-prices to remove systematic bias | Stable bias-to-surface mapping assumed; recalibration warps over months | scipy.optimize.minimize, constrained linear
regression |
| 12 | Base-rate outside-view priming | Judgmental forecasting | Kahneman & Tversky (1973), Psych. Rev. 80(3): 237–251,
doi:10.1037/h0034749 [T1] (MiniMax); Qwen
attributes to "Tetlock (2015)" with a 2010 newsletter URL that predates
it |
Force explicit reference-class construction before any probability claim is entered in the ledger | Reference-class composition drifts and its selection is arbitrary enough to introduce bias | Elicitation-UI discipline; no library required |
| 13 | Extremizing — linear blend, α ∈ [0.05, 0.15] per iteration | Judgmental forecasting | Baron, Mellers, Tetlock, Stone & Ungar (2014), cited by MiniMax
as Psychological Science 25(2): 437–444,
doi:10.1177/0956797613504262 [T2] — title/venue
flagged by MiniMax's own digest as mismatched to the extremizing
result |
Shift the private forecast p̂ away from a market consensus m toward 0 or 1 by a small fraction per update, under a Brier objective | Choice of α is fragile to regime; score is asymmetric under partial pooling. Requires holdout validation before production use | Manual linear combination |
| 14 | Extremizing — logit form, scaling exponent d = 1.4 | Judgmental forecasting | Satopää et al. (2014), Annals of Applied Statistics 8(2):
916–940, doi:10.1214/14-AOAS752; Baron et al. (2014), Decision
Analysis 11(2): 133–145, doi:10.1287/deca.2014.0293
[T2] (Gemini) |
Logit-scale an underconfident crowd or model-ensemble consensus before pricing against a contract. Not comparable to row 13's α — different operation, different scale | GJP questions were static and long-horizon; order books reprice instantly on news | scipy.optimize, NumPy |
| 15 | Trimmed-mean / geometric-mean aggregation | Judgmental forecasting | Mellers, Stone, Murray et al. (2015), Persp. Psych. Sci.
10(3): 267–281, doi:10.1177/1745691615576804 [T1]
(MiniMax) |
Aggregate several probability sources by trimmed mean rather than simple average; Brier-improving | Forecasters are non-independent and biased in different directions; bias-correct before aggregating, not after | numpy.mean on trimmed array |
| 16 | Linear opinion pooling / combining forecasts | Judgmental forecasting | Clemen (1989), Int. J. Forecasting 5(4): 559–583,
doi:10.1016/0169-2070(89)90012-8 [T1]; Cooke (1981),
Experts in Uncertainty [T2]
(MiniMax) |
Pool venue consensus, economist-survey medians (e.g. SPF), and private signals into one probability | GAP — MiniMax's body cites this without a stated
transfer risk; absent from its Table D and its consolidated risk
table |
NumPy weighted combination |
| 17 | Track-record / accuracy-weighted aggregation | Judgmental forecasting | Satopää (2014), PhD thesis [T3] (MiniMax;
institutional handle flagged as implausible by its own digest);
operational form in Mellers et al. (2015) |
Weight each probability source by its own historical accuracy rather than equally | Track-record validity under non-stationary environments is
explicitly an open question in the source [T3] |
Manual weighting; see rows 44–46 for the IRT formulation |
| 18 | Frequent updating as tournament discipline | Judgmental forecasting | Mellers et al. (2015), doi:10.1177/1745691615576804
[T1] (MiniMax, Qwen) |
Treat each macro release as a scored tournament tick and re-price immediately rather than holding a stale position | Calibration collapse under stake size: probability estimates compress toward 0.5 when the spread is non-trivial, violating proper-scoring theory | Event-driven loop; no library required |
| 19 | Prediction markets versus prediction polls | Judgmental forecasting | Atanasov, Reshetar, Zhang & Zwick (2020), cited by MiniMax as
Management Science 66(9): 4076–4094, doi:10.1287/mnsc.2019.2269
[T2] — author list, volume and DOI flagged as
implausible for the title by MiniMax's own digest |
Use venue-implied probabilities as an input to the trader's hierarchical pool rather than as a competitor to it | Endogeneity at size. Directional claim (do superforecasters beat markets?) is contested between Qwen and MiniMax and is excluded — see Resolved Conflicts | Venue API extraction; NumPy |
| 20 | Delphi method | Judgmental forecasting | Rowe & Wright (1999) [T2] (Qwen only; URL is a
course-site mirror, not a publisher host) |
Structured multi-round elicitation and aggregation of expert judgments into a consensus forecast | Groupthink and facilitator influence; Qwen further notes it is "less applicable to anonymous online markets" and that a solo retail participant cannot run it | Not applicable to a single participant |
| 21 | Verbal-to-numeric elicitation for rare events | Judgmental forecasting | Fischhoff & Davis (2014), WIREs Climate Change,
doi:10.1002/wcc.318 [T2] — topic mismatch flagged:
the cited paper is on climate-uncertainty communication
(MiniMax) |
Convert qualitative conviction into a reference-class PMF before it enters the ledger | GAP — no transfer risk stated for this row in any
source |
Elicitation UI |
| 22 | Question decomposition into tractable components (Fermi-ization) | Judgmental forecasting | Named as a core GJP practice by Qwen with no originating citation
supplied [T6] on provenance (Qwen only) |
Break a compound contract question ("will CPI exceed 3.0% AND the Fed hold?") into separately estimable components before assigning a probability | GAP — Qwen names the practice without stating a
transfer risk, and supplies no citation to check |
Elicitation-UI discipline; no library required |
| 23 | Cramér–Lundberg ruin model and Lundberg adjustment coefficient | Actuarial science | Lundberg (1903); Cramér (1930); modern treatment Asmussen &
Albrecher (2010), Ruin Probabilities [T2]
(MiniMax, Qwen, Gemini). Surplus
U(t) = u + ct − S(t),
ψ(u) = P(inf U(t) < 0) (Qwen); Lundberg
inequality ψ(u) ≤ e^(−Ru), R the unique positive root of
λ + cR = λM_X(R) (Gemini) |
Model daily P&L as a surplus process and compute P(the $100 stake is depleted before day 90) under the candidate strategy's empirical return distribution | Claim sizes assumed i.i.d. and exogenous; trading returns are serially correlated, non-stationary, and exhibit tail clustering. Ruin theory does not condition on the process changing in response to the analyst's own signal, so estimates are systematically optimistic where it does | numpy.random Monte Carlo; Lundberg exponent in ~50
lines scipy.optimize; lifelib primitives |
| 24 | Collective risk model — frequency-severity decomposition | Actuarial science | Panjer (1981), ASTIN Bulletin 12(1): 22–26,
doi:10.1017/S0515036100006615 [T1] (MiniMax,
Qwen) |
Decompose any candidate return process into trade frequency and per-trade severity before any other analysis; severity is the log-return right tail, frequency the purged effective observation count | Empirical distributions overfit against the assumed Poisson/negative-binomial family; use an empirical bootstrap as cross-check | Manual, ~30 lines |
| 25 | Panjer recursion | Actuarial science | Panjer (1981), doi:10.1017/S0515036100006615 [T1]
(MiniMax) |
Compute the exact finite-horizon aggregate P&L
distribution S = X₁ + … + X_N for small trade counts
(source specifies K = 10/quarter), avoiding asymptotic approximations
that are worthless at N = 10 |
Distribution-family mismatch between the assumed compound family and realized returns | Manual, ~30 lines |
| 26 | Bühlmann credibility, Z = n/(n+K) |
Actuarial science | Bühlmann (1967), ASTIN Bulletin 4(3): 199–207,
doi:10.1017/S0515036100008832 [T1] (MiniMax) |
Shrink a strategy-edge estimate toward the population mean of pre-registered retail strategies, with weight rising in observation count | Risk classes assumed mutually independent; candidate strategies are correlated through shared macro factors. K is unknown for a new strategy | Manual; brms/PyMC for the hierarchical
extension |
| 27 | Bühlmann–Straub credibility | Actuarial science | Bühlmann & Straub (1970), Mitt. Ver. Schweiz.
Versicherungsmathematiker 70: 111–133 [T1]
(MiniMax, Gemini); Qwen attributes to Bühlmann & Gisler
(2005) |
Multi-level credibility with explicit measurement-error structure; blends backtest alpha with retail base rates | Assumes the underlying risk process is stationary — Qwen states this is "often false in financial markets"; Gemini adds tail clustering | NumPy custom module |
| 28 | Bayesian credibility (credibility as conjugate Bayes) | Actuarial science | Jewell (1974), Geneva Papers on Risk and Insurance Theory
1(1): 77–80, doi:10.1007/BF02553258 [T2] — title
garbled in source ("Bayesian Bayesian"); year/volume pairing
flagged (MiniMax) |
Establishes Bühlmann credibility as exact Bayes under a conjugate prior, licensing direct Bayesian implementation | GAP — MiniMax's body cites it without a stated transfer
risk; absent from Table D and the consolidated risk table |
scipy.stats conjugate updates |
| 29 | Extreme-value theory — GPD peaks-over-threshold | Actuarial science | Embrechts, Klüppelberg & Mikosch (1997), Modelling Extremal
Events [T2]; McNeil, Frey & Embrechts (2015),
Quantitative Risk Management [T2] (MiniMax,
body only) |
Fit the empirical POT Generalized Pareto tail on each candidate strategy's worst 5% of observations and verify fit before trusting any estimated Sharpe | GAP — MiniMax gives an operational recipe but states no
transfer risk for this row; it is one of ~12 body-only techniques absent
from its Table D |
scipy.stats.genpareto; manual POT fit |
| 30 | Loss-development triangles / chain-ladder | Actuarial science | Mack (1993), ASTIN Bulletin 23(2): 213–225,
doi:10.1017/S0515036100009412 [T1] (MiniMax) |
None — explicitly disclaimed. Listed because a reader will encounter it in the actuarial literature | Source states: "Largely irrelevant; do not transfer"
[T6] on relevance. Retained as an explicit exclusion, not a
recommendation |
chainladder — not needed for this problem |
| 31 | Bayesian nowcasting under reporting delay | Epidemiology | Höhle & an der Heiden (2014), Biometrics 70(4):
993–1002, doi:10.1111/biom.12194; generalized in Günther et al. (2021),
Biometrical Journal 63(8): 1575–1593,
doi:10.1002/bimj.202000112 [T1] (MiniMax) |
Daily-updated estimate of a latent macro variable from sparse, delayed observations between scheduled releases | Reporting system assumed exogenous; macro revisions are strategic and bidirectional, not merely delayed. Noise model must jointly specify measurement error, seasonality and revisions or intervals are overconfident | PyMC with an explicit reporting-delay layer |
| 32 | NobBS Bayesian delay nowcasting |
Epidemiology | McGough, Johansson, Lipsitch & Menzies (2020), PLOS Comp.
Biol. 16(4): e1007735, doi:10.1371/journal.pcbi.1007735
[T1] (Gemini) |
Correct reporting delays and backfill in BLS/GDP/CPI release series to produce a current-state estimate | Clinical reporting delays are physical; economic data are strategically revised | PyMC |
| 33 | Reporting-delay decomposition and backfill correction | Epidemiology | Höhle & an der Heiden (2014); Günther et al. (2021)
[T1] (MiniMax) |
Re-estimate each past probability once later information completes, then re-score; extracts more calibration information per settled contract than one-shot scoring | Past estimates are not only systematically low but also noisy; combine with bootstrap CIs before acting on the correction | PyMC; manual re-scoring loop |
| 34 | Mixed-frequency nowcasting (MIDAS) | Epidemiology (MiniMax labels this "Macro-econometrics") | Giannone, Reichlin & Small (2008), J. Monetary Econ.
55(4): 665–676, doi:10.1016/j.jmoneco.2008.05.010 [T1]
(MiniMax) |
Kernel-weighted regression on mixed-frequency observations producing a daily probability surface over "will CPI exceed X at the next release?", priced against the corresponding contract | Mixed-frequency weighting is fragile to publication-calendar changes; macro data are conditioned on prior announcements and revisions | statsmodels; manual kernel-weighted lag regression |
| 35 | Hierarchical Bayesian partial pooling | Epidemiology | Gelman & Hill (2007), Data Analysis Using Regression and
Multilevel/Hierarchical Models [T2]; Carpenter et al.
(2017), J. Stat. Software 76(1), doi:10.18637/jss.v076.i01
[T1] (MiniMax) |
Pool probabilities across venues (Kalshi, ForecastEx, IBKR) with venue-specific intercepts and a common latent-state loading — strictly better than any single venue when venues are partially segmented | Cross-market arbitrage collapses the mispricings the pooling exploits, so value falls as venues integrate; vendor change and regime shift break the pooling structure | PyMC, cmdstanpy/Stan,
brms |
| 36 | Kalman-filter macro nowcaster (fully specified) | Epidemiology / signal processing | Kalman (1960), doi:10.1115/1.3662552 [T1]; Harvey
(1989), Forecasting, Structural Time Series Models and the Kalman
Filter [T2] (MiniMax) |
State = (latent inflation, latent unemployment, latent recession probability); observations = (released CPI, released NFP, venue-implied probabilities); transition = AR(1) latent drift. Posterior mean prices directly against contracts | State transition assumed linear-Gaussian; financial series are fat-tailed. Non-Gaussian observation noise degrades the posterior | filterpy.kalman,
statsmodels.tsa.statespace; ~100 lines NumPy |
| 37 | CUSUM change-point detection | Signal processing (MiniMax labels "Industrial statistics") | Page (1954), Biometrika 41(1/2): 100–115 [T1]
— DOI conflict: MiniMax gives 10.1093/biomet/41.1-2.100, Gemini
gives 10.2307/2333009 (MiniMax, Gemini) |
Run CUSUM on running expected log-return or risk-normalized P&L; breach of an ARL-tuned threshold declares strategy drift | In-control and post-change distributions assumed fixed; both drift continuously. ARL inflation is significant when parameters are estimated from data, so real false-alarm rates exceed nominal | ruptures;
statsmodels.stats.diagnostic.breaks_cusumolsresid.
MiniMax's ruptures.detect.cusum does not
exist |
| 38 | GLR-CUSUM | Signal processing | Lorden (1971), Ann. Math. Stat. 42(6): 1897–1908,
doi:10.1214/aoms/1177693014 [T1] (MiniMax) |
CUSUM that estimates the post-change parameter rather than fixing it — the recommended variant because post-degradation behaviour is unknown ex ante | Requires an estimate of post-change parameters; same drift problem as row 37, plus strategy detection bias (the position itself causes the drift being measured) | Manual implementation; ruptures |
| 39 | Bayesian Online Change-Point Detection (BOCPD) | Signal processing | Adams & MacKay (2007), arXiv:0710.3742 [T3];
refereed treatment Fearnhead & Liu (2007), JRSS B 69(4):
589–605, doi:10.1111/j.1467-9868.2007.00545.x [T1]
(MiniMax, Gemini) |
Posterior over run length on daily P&L with a hazard rate tuned to expected strategy half-life; P(run length > k) is instantaneous strategy credibility. Gemini applies it to order-book regime shifts for stop-out triggering | Hazard/run-length prior is fragile and concept drift produces multiple overlapping changes; assumes Gaussian white noise where returns are jump-diffusion. Mitigation: heavy-tailed run-length prior | ruptures; bayesian-changepoint-detection
(existence unverified) |
| 40 | Sequential Probability Ratio Test (SPRT) | Signal processing and clinical trials (dual-domain; Qwen assigns to industrial statistics) | Wald (1945), Ann. Math. Stat. 16(2): 117–186,
doi:10.1214/aoms/1177731118 [T1] (MiniMax, Gemini, Qwen
— Qwen's URL is an Instagram Reel and is dropped) |
Test H₀ (win rate = 50%) against H₁ (60%) at α = 0.05, β = 0.20 as a
stopping rule for a single strategy; source states termination after ~30
trades on average under a true 60% rate [T6] on the
figure |
Assumes i.i.d. observations; trade P&L is serially correlated through overnight gaps and macro cycles. Remedy is effective sample size in the threshold computation — the Bartlett/Wald §5 attribution for this is flagged as implausible. Qwen adds: binary-hypothesis design is awkward for continuous forecasts | Manual computation; ESS correction; gsDesign port for
the interim-monitoring form |
| 41 | Wald–Wolfowitz SPRT optimality | Signal processing | Wald & Wolfowitz (1948), Ann. Math. Stat. 19: 326–329,
doi:10.1214/aoms/1177699121 [T1] (MiniMax, body
only) |
Establishes that SPRT minimizes expected sample size among all tests at the same α and β — the guarantee that makes row 40 worth using at N ≈ 30, and the result underwriting rows 51–54 | GAP — MiniMax's body cites it without a stated transfer
risk; absent from its Table D |
n/a — theoretical guarantee |
| 42 | Shewhart control chart | Signal processing | Shewhart (1924), Economic Control of Manufactured Product;
Montgomery (2019), Introduction to Statistical Quality Control,
8th ed. [T2] (MiniMax) |
Lightweight three-sigma regime detection on P&L | Low statistical power against small shifts — the shifts most likely to matter at this sample size | Manual; dashboard |
| 43 | EWMA control chart | Signal processing | Roberts (1959), Technometrics 1(3): 239–250,
doi:10.1080/00401706.1959.10489860 [T1]
(MiniMax) |
Single smoothed deviation-from-target with two-sigma bands; breach declares regime change. Cheapest instrument in this section | Sensitive to the smoothing-parameter choice; a regime change produces a permanent shift the chart treats as transient. Mitigation: multiple horizons plus CUSUM as backup | Manual, ~10 lines |
| 44 | Extended / Unscented Kalman filter | Signal processing | Julier & Uhlmann (1997), Proc. AeroSense; Julier & Uhlmann
(2004), Proc. IEEE 92(3): 401–422,
doi:10.1109/JPROC.2004.823170 [T1] (MiniMax, body
only) |
Latent-state estimation where the observation or transition map is nonlinear | GAP — named in MiniMax's §19.2 table without a stated
transfer risk; absent from its Table D |
filterpy |
| 45 | Particle filter | Signal processing | Gordon, Salmond & Smith (1993), IEE Proc. F 140(2):
107–113, doi:10.1049/ip-f-2.1993.0014 [T1]
(MiniMax) |
Non-Gaussian latent-state estimation — e.g. "which of three discrete volatility regimes is active"; ~150 lines for a one-dimensional state | Computational cost is the binding constraint; mitigation is conjugate approximation or Rao-Blackwellization where feasible | filterpy.monte_carlo |
| 46 | Rasch model (1-parameter IRT) | Psychometrics | Rasch (1960), Probabilistic Models for Some Intelligence and
Attainment Tests [T2] (MiniMax, Gemini) |
Treat each historical forecast as an item and the trader's calibration as a single latent-trait parameter, estimating item difficulty jointly so that easy and hard contracts are not scored alike | Trivial-N: at N = 10–30 contracts the estimator is not identified. Source states the trader has enough data for a Bayesian IRT with strong priors, not for a Rasch fit. Gemini adds: psychometric traits are stable, trader skill is state-dependent | pyirt; manual EM |
| 47 | 2-parameter logistic IRT | Psychometrics | Birnbaum (1968), in Lord & Novick, Statistical Theories of
Mental Test Scores [T2] (MiniMax) |
Estimate each information source's discrimination and difficulty separately rather than a single accuracy number | Same trivial-N and non-stationarity problems as row 46 | pyirt |
| 48 | 3-parameter logistic IRT (adds guessing parameter) | Psychometrics | Rasch (1960) / Lord (1980) [T2]/[T3]
(Gemini) |
Separate forecaster skill θ from contract difficulty b_j and guessing c_j — the correct structure for binary contracts where a coin flip scores 50% | Psychometric traits assumed stable; trader skill fluctuates with stake, regime and fatigue | scipy.optimize custom 3PL |
| 49 | Polytomous / graded-response IRT | Psychometrics | Samejima (1969), Psychometrika 34(4): 1–97,
doi:10.1007/BF03390160 [T2] — pagination flagged as
monograph-supplement, not a regular article (MiniMax, body
only) |
Extends rows 46–48 to ordered multi-outcome contracts rather than binary settlements | GAP — named without a stated transfer risk; absent from
MiniMax's Table D |
pyirt extensions |
| 50 | Hierarchical / Bayesian IRT | Psychometrics | Fox (2010), Bayesian Item Response Modeling
[T2] (MiniMax) |
The identified alternative to row 46 at this sample size: strong priors plus population pooling in place of a free Rasch fit | Item difficulty drifts; requires time-bounded parameter estimation | PyMC |
| 51 | Empirical-Bayes (EAP) ability estimation | Psychometrics | Bock & Mislevy (1982), Applied Psych. Measurement 6(4):
431–444, doi:10.1177/014662168200600405 [T1] —
author initials flagged as transposed in source
(MiniMax, body only) |
Posterior-mean ability estimate that is stable at small N, unlike maximum likelihood | GAP — named without a stated transfer risk; absent from
MiniMax's Table D |
pyirt; manual EAP quadrature |
| 52 | Hierarchical rater model | Psychometrics | Patz, Junker, Johnson & Mariano (2002), ETS Research Report
[T5] — grey literature; a peer-reviewed version exists and
would be the better citation (MiniMax) |
Cluster signals by source with source-level and contract-level parameters; MCMC posterior over each source's reliability | Rater errors correlated through shared source bias — this model is itself the stated mitigation for row 53's violation. MCMC convergence is the practical risk | PyMC |
| 53 | Dawid–Skene latent-truth model | Psychometrics | Dawid & Skene (1979), JRSS C 28(1): 20–28,
doi:10.2307/2346806 [T1] (MiniMax) |
Treat each signal as a rater and each contract resolution as an item; EM jointly estimates latent truth and per-source error rates over a sliding window of N = 100 contracts | Assumes conditional independence of rater errors, which correlated signals (all reading the same news) violate. Non-stationary source quality biases the estimates; remedy is a time-bounded rolling re-fit | Manual EM, ~30 lines |
| 54 | Generalizability theory (G-theory) | Psychometrics | Cronbach, Gleser, Nanda & Rajaratnam (1972), The
Dependability of Behavioral Measurements [T2]
(MiniMax, body only) |
Decompose reliability into within-source, between-source and item-heterogeneity variance — separating "fragile to the choice of source" from "genuinely high signal quality" | GAP — named and operationalized in MiniMax's §20.3
without a stated transfer risk; absent from its Table D |
Variance-components estimation; statsmodels mixed
models |
| 55 | Shannon entropy and mutual information | Information theory | Shannon (1948), Bell System Technical Journal 27(3):
379–423 and 27(4): 623–656, doi:10.1002/j.1538-7305.1948.tb01338.x
[T1] (MiniMax) |
Quantify outcome uncertainty and the dependence between a candidate signal and the realized outcome; rank signals by estimated MI rather than by raw accuracy, which is blind to redundancy | Joint distribution drifts. MI estimated from finite samples
is biased upward (MiniMax §21.4); that this bias is most severe
at N ≈ 10–30 is an inference by this merge, not a source statement
[T6]. Mitigation: KSG estimator, rolling re-estimation |
dit;
sklearn.feature_selection.mutual_info_classif (biased); KSG
implementation for production use |
| 56 | Kelly criterion / log-optimal growth | Information theory | Kelly (1956), Bell System Technical Journal 35(4): 917–926,
doi:10.1002/j.1538-7305.1956.tb03809.x [T1] (MiniMax,
Gemini) |
Reference framework for sizing; identifies the achievable growth
rate with the mutual information between signal and
outcome. MiniMax's f* = I(X;Y)/H(X) is a category
error — a rate is not a capital fraction — and is excluded from this
report |
Assumes a stationary distribution and that the true probability is known; misestimating P causes catastrophic over-betting rather than mild inefficiency. Separately, Kelly optimizes long-run geometric growth, which is the wrong objective for a fixed-multiple target under a hard deadline | Manual; cvxpy for constrained sizing |
| 57 | Multi-dimensional / log-optimal portfolio Kelly | Information theory | Cover & Thomas (2006), Elements of Information Theory,
2nd ed., Ch. 6, doi:10.1002/047174882X [T2]
(MiniMax) |
Simultaneous allocation across several correlated bets | Mathematically delicate; source states it is rarely advisable without simulation. Correlated-signal structure is the binding difficulty | Numerical optimization, cvxpy |
| 58 | KL divergence as a contract screen; entropy pooling | Information theory | Kullback & Leibler (1951), Ann. Math. Stat. 22: 79–86,
doi:10.1214/aoms/1177729694 [T1]; entropy pooling in
finance, Meucci (2010), "Fully Flexible Views" [T4]
(MiniMax, Gemini) |
D_KL(p̂ ‖ m) between subjective probability and market
mid-price screens contracts for expected log-growth; entropy pooling
projects a subjective view onto the market-implied prior under arbitrary
constraints, generalizing Black-Litterman |
The growth identity holds only when p̂ is the true probability. Where p̂ merely differs from m, large divergence signals large expected loss as readily as gain — as stated in the source it licenses "bet wherever you disagree with the market." Usable only downstream of demonstrated calibration. Entropy pooling additionally fails when view constraints are jointly infeasible | cvxpy projection; scipy.stats.entropy; ~50
lines for entropy pooling |
| 59 | Maximum-entropy priors | Information theory | Jaynes (1957), Physical Review 106: 620–630,
doi:10.1103/PhysRev.106.620 [T1] (MiniMax) |
Least-committal distribution consistent with known moments, for signal sources that are genuinely model-free | GAP — named and operationalized in MiniMax's §21.3
without a stated transfer risk; absent from its Table D |
scipy.optimize under moment constraints |
| 60 | Fano's inequality | Information theory | Fano (1961), Transmission of Information, MIT Press
[T2] — cited in MiniMax's body and Table D with no
bibliographic record anywhere in the file
(MiniMax) |
Lower-bounds misclassification probability given I(X;Y), hence a lower bound on the risk of the trader's bets | Finite-sample MI bias (row 55) propagates directly into the bound, making it optimistic | Manual computation |
| 61 | Group sequential design and the Pocock boundary | Clinical trials | Pocock (1977), Biometrika 64(2): 191–199,
doi:10.1093/biomet/64.2.191 [T1] (MiniMax,
Gemini) |
Pre-specified interim analyses with equal α spent per look, so that scheduled P&L reviews do not inflate the false-positive rate | The number of looks must be fixed in advance and is itself a design parameter; a trader who peeks off-schedule invalidates the boundary. Mitigation: a self-binding software schedule | gsDesign (R) port; manual boundary tables |
| 62 | O'Brien–Fleming boundary | Clinical trials | O'Brien & Fleming (1979), Biometrics 35(3): 549–556,
doi:10.2307/2530245 [T1] (MiniMax, Gemini) |
Very conservative early boundary that spends almost no α at the first looks — matched to a setting where an early false positive is the expensive error | Same fixed-look-count requirement as row 61 | gsDesign boundaries; scipy.stats.norm |
| 63 | Lan–DeMets alpha-spending function | Clinical trials | Lan & DeMets (1983), Biometrika 70(3): 659–663,
doi:10.1093/biomet/70.3.659 [T1] (MiniMax,
Gemini) |
Continuous alpha-spending that removes the requirement to fix the number of looks in advance; controls cumulative type-I error at 0.05 across all interim reviews | Patient outcomes are independent; trade returns are serially autocorrelated, so the nominal boundary understates true spending. Same ESS remedy as row 40 | scipy.stats.norm spending function;
gsDesign/rpact port |
| 64 | Information-time versus calendar-time alpha spending | Clinical trials | Lan & DeMets (1989), Stat. in Medicine 8(10):
1191–1198, doi:10.1002/sim.4780081003 [T1]
(MiniMax) |
With ≤30 trades over 90 days, calendar fraction and information fraction diverge sharply; source concludes this problem should spend alpha on information time, since effective sample size lags the calendar | The definition of information time is itself sensitive — it requires an effective-sample-size estimate that serial correlation makes uncertain | Manual; rpact |
| 65 | Haybittle–Peto boundary | Clinical trials | Haybittle (1971); Peto et al. (1976) [T6] —
MiniMax flags its own appropriateness claim as author inference
and both citations as requiring confirmation; neither has a complete
bibliographic record (MiniMax) |
All interim looks evaluated at α ≈ 0.001, full α reserved for the final analysis — the most aggressive available protection against declaring edge early on noise | Same fixed-look-count requirement as rows 61–62 | Manual boundary |
| 66 | Pre-registration and protocol lock | Clinical trials | ClinicalTrials.gov guidance; FDA Modernization Act (1997)
[T5] (MiniMax) |
Written protocol before launch binding three components: the parameterized hypothesis, the a priori stopping rule, and the look-elsewhere correction recording how many candidate strategies were considered | Over-engineered at this capital scale — institutional pre-registration cost is amortized over millions of dollars. Source's resolution: "The transfer is at the discipline, not at the registry." Also, a solo trader can quietly revise the protocol; mitigation is a locked, timestamped document | Locked PDF; version control |
| 67 | CONSORT reporting standard | Clinical trials | Schulz, Altman & Moher (2010), BMJ 340: c332,
doi:10.1136/bmj.c332 [T1] (MiniMax) |
Reporting template ensuring the final write-up states what was pre-specified, what was changed, and what was excluded | Protocol flexibility under stress — the standard constrains reporting, not behaviour | Manual discipline |
| 68 | Bayesian sequential design (posterior predictive) | Clinical trials | Spiegelhalter, Abrams & Myles (2004), Bayesian Approaches to
Clinical Trials and Health-Care Evaluation [T2];
Jennison & Turnbull (2000) [T2] (MiniMax) |
Declare success when P(edge > 0 | data) > 0.95 under a beta-binomial conjugate posterior; source calls this the operationally cleaner alternative to alpha-spending for a single trader, and it handles unscheduled looks natively | Model misspecification — the conjugate posterior assumes a fixed win probability across a sample where the underlying rate may be drifting. Mitigation: posterior predictive checks | scipy.stats.beta; PyMC. MiniMax's
betaind "from scipy" does not exist |
Table D row count: 68. Techniques carrying a
GAP marker (named by a source with no transfer risk stated
anywhere in that source): rows 2, 16, 21, 22, 28, 29, 41, 44, 49, 51,
54, 59 — 12 rows. Ten of the twelve are techniques
MiniMax treats in its §15–§22 body but omits from its own Table D; the
merge preserves them rather than dropping them, at the cost of an
unstated risk cell. Row 22 is Qwen-only. Row 30 (chain-ladder) is
not a gap — it is a deliberate non-transfer with an
explicit exclusion stated by its source, and is excluded from this
count.
[T2] vs [T3]); MiniMax's thesis retained on
the separate track-record-aggregation row with its [T3]
grading and the institutional flag intact.f* = I(X;Y)/H(X).
MiniMax states this, attributes it to Cover & Thomas, and builds its
signal-selection prescription on it. Its own digest identifies it as a
category error — Kelly's identity equates a growth rate with
mutual information, whereas a bet fraction is a dimensionless capital
share. Gemini states the correct form (growth rate ↔︎ channel capacity)
without the sizing claim. Resolved: formula excluded
entirely from this report. The correct rate identity and the
signal-ranking use survive; the sizing rule does not.NobBS; MiniMax gives
Höhle & an der Heiden (2014) generalized by Günther et al. (2021).
Resolved: not a conflict — Höhle & an der Heiden is
the originating Bayesian nowcasting method and McGough et al. the widely
used implementation. Both carried on adjacent rows with roles
distinguished.[T6].[T6]. No replacement figure
invented, per the source's own flagging convention.properscoring — recommended and blacklisted
simultaneously. Gemini's AVOID list marks it unmaintained since
2015-05-20 while Gemini's own Table D recommends it; MiniMax recommends
it and dates the last release to 2017, flagging the version
[T6]. Resolved: scores
recommended as primary; properscoring retained as a
secondary with the maintenance status flagged and the two conflicting
release dates both recorded as unverified. MiniMax's observation that
Brier and CRPS are under 30 lines of NumPy each makes the dependency
avoidable.betaind "from scipy" (§22.5) and
ruptures.detect.cusum (Table D row 21); neither exists —
ruptures exposes search classes (Pelt,
Binseg, Window, BottomUp,
Dynp) with cost functions, and the correct SciPy symbol is
scipy.stats.beta, which MiniMax's own Table D uses.
Resolved: both corrected to the forms MiniMax itself
gives elsewhere, with the fabricated symbols named explicitly so
downstream readers do not reintroduce them.[T1] and its digest instructs that T1 counts not be
aggregated across reports; MiniMax graded a NOAA technical bulletin and
ten monographs [T1]. Resolved: all tags in
this section re-derived under the convention stated in §6.0. Monographs
and textbooks (Cover & Thomas, Asmussen & Albrecher, Gelman
& Hill, Montgomery, Rasch 1960, Fox 2010, Cronbach et al. 1972)
placed at [T2]; grey literature (Patz et al. ETS report,
ClinicalTrials.gov) at [T5]; preprints and theses (Adams
& MacKay arXiv, Satopää thesis) at [T3];
author-constructed numerics (the 1–5% base rate, the ~30-trade SPRT
figure, the 0.49 alpha inflation, all package versions) at
[T6].Coverage gap across sources. Qwen covers four of the eight required source domains — meteorology, judgmental forecasting/GJP, actuarial science, and SPRT under an industrial-statistics label. It omits epidemiology, signal processing as a domain, psychometrics, information theory, and clinical-trial methodology entirely. Gemini and MiniMax each cover all eight. Every technique in §6.4, §6.6, §6.7, and §6.8 therefore rests on at most two independent reports, and roughly two-thirds of Table D's rows in those domains rest on MiniMax alone. Sections 6.1, 6.2, and 6.3 are the only parts of this section with three-report support.
This section carries the same weight as Section 5. That is a deliberate design choice, not a courtesy. For a USD 100 stake over 90 days, the set of strategies that reliably destroy capital is larger, better documented, and more consequential to the outcome than the set that might preserve it. A retail participant who correctly excludes the nine categories below has done more to protect the experiment than one who correctly identifies a marginal edge, because the exclusions are near-certain and the edges are not. All three source reports converge on this asymmetry independently (Source: Qwen, Gemini, MiniMax).
We grade each category on four axes: whether a genuine edge exists after costs, whether the input signal is a structural feature of the return-generating process or an artifact of the search that found it, whether the sizing rule can survive its own worst sequence, and whether the result replicates out of sample under multiple-testing correction. MiniMax's cross-cutting synthesis names these four pathologies explicitly — no edge, no signal, no risk control, no replication — and they organize the section cleanly (Source: MiniMax).
One caution about what follows. Several of the numbers in the source reports did not survive verification, and we say so where that happened rather than laundering them into the master report. Where the three reports disagree on a figure, we report the better-sourced value and log the conflict at the end of this section. Where all three rest on a citation we could not place, we drop the number and keep only the argument. The verdicts in this section are robust; a minority of the supporting figures are not, and we mark the difference.
Two tagging conventions apply throughout.
[T1, derived] marks arithmetic that
applies a peer-reviewed formula to explicitly stated parameters: the
formula carries T1 support, the computation is ours, and every input is
printed so the reader can recompute it. This is distinct from
[T6], which marks a figure asserted by a source or by us
without a reproducible derivation. Second, where a source supplied a DOI
that its own digest flagged as inconsistent with the stated journal, we
reproduce the identifier as given, with the flag
attached, rather than silently omitting it or inventing a
replacement — an unverified identifier the reader can check is more
useful than no identifier at all, provided it is labeled.
Verdict: negative expected value after costs, with an unusually clean historical explanation for why the early evidence looked positive.
The technical-analysis literature is not a story of a claim that was never supported. It is a story of a claim that was supported on pre-1988 data and then stopped being supported, and the distinction matters because it explains why practitioner belief persists.
Brock, Lakonishok and LeBaron (1992) reported that moving-average and trading-range-breakout rules generated statistically significant returns on the Dow Jones Industrial Average across roughly a century of daily data [T2] (DOI 10.1111/j.1540-6261.1992.tb04681.x). That paper is the canonical affirmative result and remains among the most-cited in the field (Source: MiniMax). Sullivan, Timmermann and White (1999) then applied White's Reality Check bootstrap to a universe of 7,846 technical trading rules on the same data and found that the Brock-Lakonishok-LeBaron results survived the data-snooping correction, though only marginally and only for a small subset of rules, with the overwhelming bulk of the 7,846 producing nothing [T2] (DOI as given by MiniMax: 10.1016/S0304-405X(99)00022-4; MiniMax names the venue as Journal of Finance while supplying a Journal of Financial Economics prefix — verify before publication) (Source: MiniMax).
MiniMax's own digest treats these two facts as contradicting its headline verdict that nothing survives multiple-testing correction. They do not contradict it once the time dimension is restored, and Gemini supplies the reconciling mechanism. Park and Irwin (2007), surveying more than 100 modern studies, found that technical profitability was real in the pre-1988 record and vanished thereafter, and attributed the disappearance to institutional algorithmic arbitrage competing the signal away [T1] ("What Do We Know About the Profitability of Technical Analysis?", Journal of Economic Surveys 21(4), 786–826, DOI 10.1111/j.1467-6419.2007.00519.x) (Source: Gemini). The correct synthesis is therefore temporal: a subset of rules survived correction on data ending in 1986, and no rule survives on data that includes the modern electronic market. Both source claims are true of different samples.
Bajgrowicz and Scaillet (2012) close the case with a false-discovery-rate correction applied to a very large universe of technical rules on a century of daily Dow data, and report that after transaction costs of 5 to 10 basis points, zero rules generate significant out-of-sample excess returns [T1] (Journal of Financial Economics 106(3), 473–491) (Source: Gemini, MiniMax). The two reports disagree on the size of the rule universe — Gemini says more than 15,000, MiniMax says 5,580 — and supply incompatible DOIs for the same paper. We report the finding, which both agree on, and log the discrepancy rather than picking a rule count neither can substantiate.
Two affirmative results deserve to be characterized precisely rather than dismissed. Lo, Mamaysky and Wang (2000) showed with Gaussian-kernel estimators that head-and-shoulders, double-bottom and triangle formations carry statistically significant conditional return differentials [T2] (DOI as given by MiniMax: 10.1016/S0304-405X(00)00065-6, attributed to Journal of Financial Economics*; MiniMax's own digest believes the paper appeared in the* Journal of Finance — verify). That is a computational-existence result: it establishes that the patterns are detectable and non-random, not that trading them survives transaction costs or replicates out of sample (Source: MiniMax). Marshall, Cahan and Young (2008) found candlestick reversal patterns significant on the Tokyo Stock Exchange and not significant on Dow and S&P 500 data under comparable methodology [T2] (DOI as given by MiniMax: 10.1093/jjfinec/nbn023, attributed to Journal of Financial Econometrics*; MiniMax's digest associates these authors' candlestick work with the* Journal of Banking & Finance instead — verify) — the standard cross-market replication failure, and a useful reminder that a single positive market is the expected output of searching several (Source: MiniMax). Both identifiers above are reproduced as the source gave them; both are flagged for venue mismatch and neither should reach publication unchecked.
Qwen adds a behavioral channel the other two omit. Retail investors who spend disproportionate time reviewing price charts exhibit significantly worse trading performance, which is consistent either with the patterns being useless or with chart-watching being a proxy for the overtrading that Section 7.2 quantifies [T2] (Barber & Odean 2000, DOI 10.1111/0022-1082.00223 — Qwen asserts the claim but sources it to a Medium post standing in for the primary literature; we reattach it to the paper Qwen was paraphrasing) (Source: Qwen). Qwen names head-and-shoulders formations and double-tops specifically as tested and failed (Source: Qwen).
Superforecaster posture. Conditional on a technical pattern having been published in a peer-reviewed venue, we put P(positive after-cost expected value in a live 2026 retail account) at roughly 5–10% [T6]. Conditional on that pattern additionally failing an independent out-of-sample replication — which is the modal outcome — the posterior falls to 0–3% [T6]. These are MiniMax's numbers and they are author-modeled, not derived from any cited study; we carry them labeled rather than dressed as findings (Source: MiniMax).
At USD 100 over 90 days. The literature tests these rules at monthly-to-annual horizons. Ninety days is too short to amortize the false-signal rate, and any pattern strategy requiring more than roughly ten trades to mature cannot clear a multiple-testing-corrected bar within the window [T6] (Source: MiniMax). The category is non-viable.
Verdict: negative expected value for the median participant, with the largest and cleanest evidence base in this entire section.
Section 3 establishes the general retail base rate. This subsection adds what is specific to day trading at USD 100 scale: the friction arithmetic, the settlement mechanics that cap turnover, and the statistical-power result that makes a 90-day day-trading experiment uninterpretable even if it succeeds.
The base-rate evidence. Barber and Odean (2000) analyzed 66,465 U.S. retail households from 1991 to 1996 and found that the average household underperformed the market — roughly 16.4% against 17.9% annually — with the highest-turnover quintile underperforming by approximately 6.5 percentage points per year [T1] ("Trading Is Hazardous to Your Wealth," Journal of Finance 55(2), 773–806, DOI 10.1111/0022-1082.00223) (Source: Gemini, MiniMax, Qwen). We state this carefully because MiniMax renders it as "the median household lost money after costs" and applies the 6.5-point gap to the median active trader. Both renderings overstate the paper: 1991–1996 was a bull market in which the average household made money while losing to the index, and the 6.5-point figure belongs to the top turnover quintile, not the median. The corrected statement is weaker and correct. Barber, Lee, Liu and Odean (2009) extend the result, quantifying the aggregate wealth transfer from individual investors through trading [T1] (Review of Financial Studies 22(2), 609–632, DOI 10.1093/rfs/hhn046).
Barber, Lee, Liu and Odean (2014) examined the entire population of Taiwanese day traders over fifteen years — approximately 1.4 million accounts on the Taiwan Stock Exchange — and found that fewer than 1% show predictable, persistent profitability net of fees [T1] (Source: Gemini, MiniMax). Gemini's digest renders this as "99% net unprofitable" and MiniMax renders it as "the top 0.1% earned positive returns; the remaining 99.9% lost money." Neither is the paper's actual finding, which concerns the fraction exhibiting repeatable skill, a strictly stronger and different claim than the fraction losing money in any given period. MiniMax additionally contradicts itself, quoting 0.1% in one paragraph and "approximately 1%" eight lines later. We report the paper's finding — under 1% with demonstrable persistent skill — and treat the 99% and 99.9% variants as unsupported amplifications.
Chague, De-Losso and Giovannetti (2020) studied 19,642 Brazilian equity-futures day traders who persisted for at least 300 trading days and found that 97% lost money, only 1.1% earned more than the Brazilian minimum wage (approximately USD 54 per day), and only 0.1% earned more than USD 300 per day [T1] (Source: Gemini). This is the most rigorous corroboration of the Taiwan result because it conditions on persistence — these are not dabblers, they are people who showed up for more than a year. MiniMax reports the same study with a different benchmark (1.1% beating Brazil's CDI overnight rate gross, 0.4% net); its own digest flags this as a re-benchmarking of the minimum-wage figures onto a different comparator. We use Gemini's framing and log the conflict.
The FINRA Investor Education Foundation reports that approximately 70% of retail forex traders lose money [T4] — a lower failure rate than equity day trading, but on an instrument with higher embedded leverage (Source: MiniMax).
The mechanism. Chague and colleagues attribute the losses to disposition bias compounded by paying bid-ask spreads to institutional market makers on every round trip (Source: Gemini). This is the correct causal story and it is not a story about being wrong more often than right. A trader who is right 50% of the time still loses at a rate set by the spread multiplied by turnover.
The USD 100 friction arithmetic. MiniMax stipulates 0.5% round-trip friction and derives a 10% drag over twenty round trips in 90 days (Source: MiniMax). That input is asserted rather than sourced, and in a zero-commission U.S. retail environment it is too high for liquid equities and too low for options. Gemini's vehicle-level figures are better grounded: 0.02%–0.10% total round-trip friction on commission-free fractional equities, 2.50%–12.00% on listed options, and 0.80%–2.50% on spot crypto (Source: Gemini). The honest range is therefore vehicle-dependent by two orders of magnitude, and the day-trading drag on USD 100 is negligible in fractional equities and ruinous in options — which is exactly the wrong way round from where retail day-trading volume concentrates. Bryzgalova and co-authors (2023) estimate retail long-premium directional options trades at −15% to −30% expected value per trade [T3, working paper] (Source: Gemini).
The settlement cap. A USD 100 account must be a cash account, because FINRA Rule 4210(f)(8)(B) requires USD 25,000 minimum equity for a margin account flagged as a pattern day trader [T5] (Source: Gemini). In a cash account, U.S. equities settle T+1 under 17 CFR § 240.15c6-1 [T5], which caps daily deployable turnover at the account balance and imposes a mandatory overnight liquidity hold (Source: Gemini). Qwen asserts T+2 three separate times and builds its turnover analysis on it; U.S. equities moved to T+1 in May 2024 and the T+2 figure is stale for a 2026 report (Source: Qwen — corrected). Three good-faith violations trigger a 90-day account restriction under Reg T [T5], which for this experiment means a single settlement mistake ends the 90-day window (Source: Gemini, Qwen).
The statistical-power result — the day-trading-specific point that matters most. To demonstrate a t-statistic of 3.0 over N = 90 trading days requires a daily Sharpe ratio of 3/√90 ≈ 0.3162, which annualizes to 5.02 [T1, derived] (Source: Gemini). Annualized Sharpe ratios above 5 essentially do not exist in unleveraged retail-accessible asset classes. A 90-day day-trading experiment therefore cannot in principle produce statistical proof of skill, regardless of outcome. Qwen reaches the same conclusion independently and states it plainly: any live result from this experiment is statistical noise, and 90 days is far too short to distinguish skill from luck in any active strategy (Source: Qwen). This is the strongest single argument against day trading in the USD 100 context, and it is stronger than the base rates because it holds even for a participant who wins.
At USD 100 over 90 days. MiniMax puts P(USD 100 → USD 200 via retail day trading) below 1% and P(ruin) at 60–70% [T6]; Gemini puts P(ruin) above 0.98 for HFT-style equity day trading or options buying [T6] (Source: Gemini, MiniMax). Neither figure is derived from a stated model. Both are author inferences and both point the same direction. The category is non-viable, and the reason is not that the trader will be wrong — it is that the trader cannot be right often enough, fast enough, to overcome friction inside a window too short to measure anything.
Verdict: deterministic negative drift relative to the leveraged benchmark, but — and this is the part both source reports get wrong — leverage genuinely raises the probability of hitting a fixed doubling target. The instrument is seductive for a real reason, and it still does not work.
The mechanism. Leveraged and inverse ETFs reset their exposure daily. Over multiple sessions the compounded result is path-dependent rather than proportional. Gemini gives the standard closed form: for a fund with leverage multiple L on an underlying with terminal price S_t and volatility σ,
X_t = X_0 · (S_t / S_0)^L · exp(½(L − L²) σ² t)
The exponential term is the volatility drag [T1] (Source: Gemini). It is negative for every L outside the interval [0, 1], which includes every leveraged long fund and every inverse fund. Avellaneda and Zhang (2010) formalize the path-dependence result in the diffusion setting [T1] (SIAM Journal on Financial Mathematics, DOI 10.1137/090771333) (Source: MiniMax). Cheng and Madhavan (2009) supply the practitioner reference and the observation that expected deviation from the daily benchmark grows with holding period [T4] (Source: Gemini, MiniMax).
The magnitudes. Gemini's worked example: for L = 3 on a flat index with σ = 25% annualized, the fund loses exp(−3 · 0.25² · 0.25) − 1 ≈ −4.6% over 90 days purely from path volatility, with no directional move at all (Source: Gemini). For L = 2 the drag term is −½σ²t per unit time. At σ = 20% annualized over 90 days, that is approximately −1.0%; at σ = 30%, approximately −2.3% [T1, derived].
MiniMax reports substantially larger figures — a −2σ² per day drift penalty for a 2× fund, yielding 2.9% over 90 days at 20% volatility and 6.5% at 30% (Source: MiniMax). MiniMax's own digest identifies the error: the standard beta-slippage drag is −(L² − L)σ²/2, which for L = 2 is −σ²/2 per unit time, not −2σ². Every downstream number in MiniMax's section is therefore roughly a factor of two to four too large. We use Gemini's formula, whose arithmetic was independently verified in digest, and drop MiniMax's magnitudes. We also drop Trainor's (2010) claimed 25–75% annual underperformance across 195 leveraged ETFs (Source: MiniMax, Gemini): it is irreconcilable with the ~2–8% annual drag the shared formula implies for broad-index funds at ordinary volatility, and MiniMax's digest flags the internal contradiction without resolving it.
The SEC's investor bulletin states the regulatory position: over time the cumulative percentage change in a leveraged ETF's NAV will likely diverge significantly from the cumulative percentage change in the underlying index [T5] (Source: MiniMax). Qwen states the mechanism qualitatively — path dependence and expense ratios produce volatility decay — but supplies no math and no figures (Source: Qwen).
The error we are not carrying forward. MiniMax concludes that P(USD 100 → USD 200 via a 90-day leveraged-ETF holding) is "essentially the same as holding the underlying at 1× leverage, less the volatility drag" (Source: MiniMax). This is analytically wrong and it is wrong in the direction that matters most for this mandate. At 1× leverage the underlying must appreciate approximately 100% to double USD 100. At 2× it must appreciate approximately 41%, since 1.41² ≈ 2. Leverage raises the probability of hitting a fixed multiplicative threshold precisely while it lowers expected value; those are different quantities and the volatility drag does not close the gap. Confusing them collapses the entire first-passage framing the report is built on.
The honest verdict is therefore more interesting than either source states it. A leveraged ETF held for 90 days is a negative-expected-value instrument that nonetheless increases P(reaching USD 200) relative to the unleveraged underlying. It is a variance purchase, and under the Dubins-Savage logic that governs a subfair fixed-target problem, buying variance is not automatically irrational. What kills it is the combination: the drag is deterministic and always adverse, the fund charges an expense ratio on top, and the same variance is available through instruments with bounded downside and no daily-reset penalty. MiniMax's own framing of leveraged ETFs as "financing cost disguised as leverage" is apt, and its conclusion that leverage at this scale is better obtained through options — where maximum loss is bounded — or not at all, survives the correction (Source: MiniMax).
At USD 100 over 90 days. Non-viable as a holding, for cost rather than probability reasons. The sign of the drag is unambiguous; the magnitude is small in absolute terms (roughly 1–5% over the window depending on L and σ) but it is a guaranteed adverse term applied to a strategy with no compensating edge.
Verdict: structurally negative expected value, driven by spread capture and manipulation rather than by directional risk. Qwen omits this category entirely.
This is the one category where the failure mode is not statistical. It is a transfer.
The spread. Bradley and co-authors (2014) studied more than 1,000 microcap and OTC issues and documented bid-ask spreads of 10% to 50% of share price, toxic convertible death-spiral dilution, pervasive pump-and-dump activity, and long-term returns approaching −100% [T1] ("Penny Stock IPOs," Journal of Banking & Finance 43, 62–73, DOI 10.1016/j.jbankfin.2014.03.003) (Source: Gemini). A 10% spread means a position must appreciate 11% before the holder breaks even on a round trip. A 50% spread means it must double. At the upper end of that range, the instrument requires the investor's target return simply to exit at cost.
MiniMax cites a competing figure — a 7–12% bid-ask markup on pink-sheet equities from Li and Zheng (2020) (Source: MiniMax). Its own digest could not place the paper and flags it as possibly fabricated. We drop it and use Bradley, whose range subsumes it anyway.
The manipulation. Aggarwal and Wu analyzed SEC enforcement actions and found pump-and-dump activity concentrated in micro-cap and OTC Bulletin Board issues, with the median manipulated stock rising 30–50% and then collapsing 60–90% within weeks [T2] (DOI as given by MiniMax: 10.1016/S0304-405X(03)00116-0, attributed to "Stock Market Manipulation and Short Selling," Journal of Financial Economics 2003; MiniMax's digest believes the actual paper is "Stock Market Manipulations," Journal of Business 2006 — verify) (Source: MiniMax). Comerton-Forde and Putniņš, using surveillance data from 35 markets, estimated manipulation prevalence and found OTC and small-cap issues significantly overrepresented [T2] (DOI as given by MiniMax: 10.1016/j.jfineco.2013.10.008, attributed to Journal of Financial Economics*; MiniMax's digest believes the venue is* Review of Finance — verify) (Source: MiniMax). We report both qualitatively. MiniMax's digest flags the year, title and venue of the first and the venue and prevalence estimate of the second as probably wrong, so the specific figures — 30–50%, 60–90%, 2.5% of trading days — should be treated as unverified rather than quoted as findings.
The regulatory signal. FINRA Rule 6432 requires broker-dealers to disclose compensation received on retail penny-stock transactions and to supply bid-ask pricing information [T5] (Source: MiniMax). MiniMax's argument from the rule's existence is sound: a disclosure regime is built where the asymmetry is severe enough to require one. We drop MiniMax's supporting claims that penny stocks represent 31% of SEC-investigated fraud dollar value and that the SEC suspended approximately 270 issuers over a 12-month period — both rest on a garbled reference to a nonexistent body ("the Securities Enforcement Commission") and a generic landing-page URL (Source: MiniMax — dropped).
Why the tail does not rescue it. The upside case for penny stocks is entering a pump early and exiting before the dump. That is the documented mechanism by which retail participants lose in this venue, not the mechanism by which they win: the coordinated operators control the timing and the retail flow is the exit liquidity. The distribution is not merely adverse in expectation, it is adverse conditional on the scenario the buyer is hoping for.
At USD 100 over 90 days. MiniMax bounds P(USD 100 → USD 200) above at 1–2% [T6] (Source: MiniMax). That figure is unmodeled. The structural observation is more useful than the estimate: with a 10–50% round-trip spread and documented manipulation concentration, this asset class transfers capital from retail to dealers and operators as a matter of market structure, independent of directional skill. Non-viable.
Verdict: no replicated positive expected value; the entire affirmative literature rests on a single event window, and the sentiment feature space is the cleanest available example of a data-snooping trap.
The sequencing problem. Nofsinger, Sault and Shank (2021) found that social sentiment metrics lag price action — retail buys at peak sentiment precisely as institutional shorting and mean reversion begin [T2] (Journal of Behavioral Finance 22(4), 412–428, DOI 10.1080/15427560.2021.1963232) (Source: Gemini). Gemini's digest could not confirm this publication's details, so we grade it T2 rather than T1. Qwen reaches the same conclusion without citation: sentiment-chasing "encourages chasing trends after they have already begun, leading to buying high and selling low," and meme dynamics constitute herding behavior devoid of fundamental analysis (Source: Qwen). This is the operative failure. The signal is not absent; it is late.
The academic affirmative results, correctly sized. Da, Engelberg and Gao (2011) showed that Google search volume predicts abnormal returns in the cross-section [T1] ("In Search of Attention," Journal of Finance, DOI 10.1111/j.1540-6261.2010.01629.x) (Source: MiniMax). MiniMax attaches a point estimate of 0.22% per standard deviation of abnormal search volume; its digest flags this as unverified against the paper, so we carry the direction and not the number [T6 for the magnitude]. Renault (2017) found StockTwits sentiment predicts intraday returns, with a per-observation magnitude small enough that economic value is contested once transaction costs net out [T2] (DOI as given by MiniMax: 10.1016/j.jfineco.2017.02.014, attributed to Journal of Financial Economics*; MiniMax's digest believes the venue is* Journal of Banking & Finance — verify) (Source: MiniMax). Bartov, Faurel and Mohan (2017) found Twitter content carries predictive power for firm-level earnings and returns over a short 2010–2015 sample sensitive to window choice [T2] (DOI as given by MiniMax: 10.1016/j.jfineco.2017.05.007, attributed to Journal of Financial Economics*; MiniMax's digest believes the venue is* The Accounting Review — verify) (Source: MiniMax). Both identifiers are reproduced as given and both are flagged for venue mismatch.
The meme episode. Pedersen (2022) analyzed the GameStop event and concluded the price spike was driven by extreme order-flow imbalance from small retail buyers rather than by any change in fundamental value or stochastic discount factors, with the imbalance exhausting itself by mid-January 2021 [T2] (DOI as given by MiniMax: 10.1016/j.jfineco.2022.07.004, titled "GameStop and the Reemergence of the Retail Investor"; MiniMax's digest believes Pedersen's 2022 JFE paper on this topic is titled "Game on: Social networks and markets," so the DOI may resolve to a different article than the one described — verify) (Source: MiniMax). Qwen names GameStop and AMC as canonical examples of extreme volatility driven by collective emotion, unpredictable and prone to violent reversals (Source: Qwen). MiniMax cites a post-2021 abnormal-return figure of −8.6% over a 36-month window from a working paper its own digest flags as highest-suspicion and absent from the file's own bibliography; we drop that citation and its number entirely (Source: MiniMax — dropped).
Why this cannot be validated even in principle. MiniMax makes the sharpest argument in this subsection and it deserves to be stated at full strength (Source: MiniMax). First, the entire meme-equity literature is contingent on one event window in January 2021 on a specific set of retail platforms. There is no out-of-sample replication across a second attention shock, because there has not been a comparable second attention shock. A strategy with one observation cannot be validated at any confidence level. Second, the feature space is enormous — text polarity, emoji counts, hashtag frequency, follower counts, retweet velocity, subreddit post volume, and arbitrary combinations and lags of each — which means the number of testable sentiment signals is effectively unbounded. Under any honest multiple-testing correction the expected value of the best-performing discovered signal converges to zero. This is Section 7.9's mechanism applied to a feature space that is unusually easy to expand, and it is why sentiment strategies backtest well and trade badly.
At USD 100 over 90 days. No demonstrated positive expected value over the window [T6]. The order-flow imbalance that produced the one profitable episode was historically specific and was not forecastable ex ante from any social-media extraction procedure (Source: MiniMax). Non-viable.
Verdict: guaranteed to overstate out-of-sample performance, by a mechanism that is well understood and fully diagnosable. The failure is not that ML does not work on markets; it is that the standard validation toolchain is invalid on this data class.
This subsection explains why, because the "that" is not in dispute and the remedies belong in Section 8.
Mechanism 1 — label overlap creates train-test leakage. Financial ML labels are typically forward returns over a horizon h: the label for observation t is a function of prices in the interval [t, t+h]. Adjacent observations therefore share outcome information. Standard k-fold cross-validation assumes exchangeability between training and test instances, and that assumption fails whenever labels overlap: an observation in the test fold shares its realized future with observations in the training fold. The model does not learn a predictive relationship, it memorizes a shared outcome. López de Prado (2018) formalizes this and the purged/embargoed corrections for it [T4, book — Advances in Financial Machine Learning, Wiley] (Source: Gemini, MiniMax). Gemini's digest records the characteristic signature: in-sample Sharpe above 4.0 collapsing to zero or negative in live trading (Source: Gemini).
Mechanism 2 — serial correlation inflates effective sample size. Financial series exhibit volatility clustering and long memory in absolute returns. Cont (2001) catalogues the stylized facts any model must respect: heavy tails, volatility clustering, long memory in absolute returns, leverage effects, and near-absence of serial correlation in raw returns [T1] (Quantitative Finance, DOI 10.1080/713665670) (Source: MiniMax). A model whose backtest violates these properties is producing artifacts. More subtly, the number of independent observations in a daily series is far smaller than the number of rows, so nominal significance thresholds computed on row counts are systematically too permissive.
Mechanism 3 — non-stationarity means the fitted relationship has no reason to persist. Qwen makes this point cleanly: financial time series are non-stationary, so their mean and variance change over time, and a model fitted to one regime is estimating a transient artifact rather than a causal law (Source: Qwen). Qwen's transfer-risk passage generalizes it well — weather evolves according to physical laws indifferent to the forecast, whereas markets are populated by agents actively seeking to exploit any predictable pattern, so the data-generating process reacts to being modeled (Source: Qwen).
Mechanism 4 — trial-count inflation. Every architecture, feature set, lookback window and hyperparameter grid searched is a trial, and the reported Sharpe is the maximum over trials. Bailey, Borwein, López de Prado and Zhu formalize the Probability of Backtest Overfitting as a function of trial count and the in-sample Sharpe distribution, and show that PBO exceeds 50% in realistic financial datasets at trial counts a single practitioner can reach in an afternoon [T1] (Source: MiniMax). Bailey and López de Prado's Deflated Sharpe Ratio supplies the corresponding correction: the probability that an in-sample Sharpe exceeds a threshold, adjusted for the number of trials, is dramatically lower than the uncorrected p-value implies [T1] (Source: MiniMax).
The scale of the correction, from the anomaly literature. Harvey, Liu and Zhu (2016) evaluated 315 published factors and concluded that the conventional t > 2.0 threshold produces widespread false discovery; the corrected threshold for a new factor claim is t > 3.0 [T1] (Review of Financial Studies 29(1), 5–68, DOI 10.1093/rfs/hhv059) (Source: Gemini, MiniMax). MiniMax renders this quotation inverted — as "t-statistics greater than 3.0 are not adequate" — which contradicts its own next sentence; we use Gemini's correct framing. Hou, Xue and Zhang (2020) replicated 452 anomalies under the q-factor model and found 65%+ fail replication at t ≥ 1.96 and 82%+ fail at t ≥ 2.78, with the failures driven by equal-weighted microcap portfolios that transaction costs destroy [T1] (RFS 33(5), 2019–2133, DOI 10.1093/rfs/hhy131) (Source: Gemini). MiniMax reports 23% surviving and 12% surviving a stricter battery; its digest flags these as not matching the paper, and Gemini's figures are the standard ones. McLean and Pontiff (2016) found published factor returns decay 58% post-publication, decomposed as approximately 26% academic overfitting plus 32% arbitrage exploitation [T1] (Journal of Finance 71(1), 5–32, DOI 10.1111/jofi.12365) (Source: Gemini). MiniMax converts these decay percentages into annual return levels — 26%/year in-sample decaying to 4%/year — which is a category error and implausible on its face; we use Gemini's correct rendering (Source: MiniMax — corrected).
What the correction implies at retail scale. MiniMax estimates that in-sample Sharpe overstates out-of-sample Sharpe by a factor of roughly 2× to 5× for a moderately naive backtest, and that a backtested Sharpe of 0.5 on five years of daily data without purged validation is indistinguishable from zero at 90% confidence, with a true out-of-sample central estimate near 0.1–0.2 [T6] (Source: MiniMax). These are unsourced author estimates and we carry them labeled. Their direction is corroborated by the replication figures above, which is why we keep them at all.
At USD 100 over 90 days. Non-viable unless wrapped in purged cross-validation with embargo, walk-forward validation, deflated-Sharpe adjustment and a multiple-testing-corrected threshold — which is Section 8's subject. Even then, Section 7.2's power result binds: 90 days cannot validate the resulting model.
Verdict: negative expected value for the subscriber, argued from equilibrium and from the fund-persistence literature by analogy. This is the weakest-evidenced category in the section, and we say so rather than manufacture support.
The honest position on evidence. No source report supplies a usable T1 or T2 citation bearing directly on retail copy-trading or paid signal services. Gemini's JSON grades the category T4 with a citation reading "Empirical Market Microstructure Analysis" — a topic label, not a work (Source: Gemini). MiniMax cites a working paper on MQL5 signals and a set of named enforcement actions — "SEC v. Managed Funds Association," "SEC v. TODAQ," a CFTC case against "CZBCrypto" — all of which its own digest flags as unverifiable or fabricated (Source: MiniMax — dropped). Qwen states the conclusion with no citation at all (Source: Qwen). We drop every unverifiable citation and rest the verdict on two things that need no citation and one literature that does.
The equilibrium argument. A signal service cannot persistently deliver positive after-cost alpha to retail subscribers, for a reason that does not depend on any empirical finding. If the provider has genuine skill, the profit-maximizing deployment of that skill is proprietary capital, and any published signal is a marketing artifact rather than the provider's actual best use of the information. If the provider lacks skill, the signal is noise sold at a price. In the intermediate case where the provider has skill and sells it anyway, subscriber flow degrades the very signal being sold, because subscribers execute after the provider and into the price impact the aggregate subscription creates. The migration of the business onto registered broker platforms — ZuluTrade, NAGA, MetaTrader signal marketplaces — changes the packaging and not the equilibrium (Source: MiniMax; note that MiniMax additionally lists "eRank," an Etsy analytics product, as a copy-trading platform — a category error we exclude).
The microstructure argument. Gemini adds the execution channel: adverse selection, execution latency, and front-running by the provider, who fills first and broadcasts seconds or minutes later (Source: Gemini). Any signal with genuine short-horizon content is worth less to the subscriber than to the provider by exactly the latency, and short-horizon content is what these services predominantly sell.
The analogy from the fund literature. The nearest well-evidenced comparator is professional manager persistence, and it is unfavorable. Carhart (1997) showed that apparent mutual-fund performance persistence is largely explained by momentum-factor exposure rather than manager skill, and that positive residual alpha in one year does not predict positive residual alpha in the next [T1] (Journal of Finance, DOI 10.1111/j.1540-6261.1997.tb03808.x) (Source: MiniMax). Fama and French (2010) found the cross-sectional distribution of mutual-fund alpha centered near zero with a thin positive tail attributable substantially to luck once multiple testing is accounted for [T1] ("Luck Versus Skill in the Cross-Section of Mutual Fund Returns," Journal of Finance 2010 — no DOI supplied by any source report; MiniMax's bibliography lists the entry without an identifier, and we do not invent one) (Source: MiniMax). If persistent skill is scarce among regulated, disclosed, professionally-resourced managers, the prior for an undisclosed retail signal seller with no audited track record and a direct incentive to inflate reported results should be lower, not higher. DALBAR's annual investor-behavior study finds the median self-directed retail investor trailing the index by 1–3% per year over 5- to 20-year windows [T4, practitioner grey literature] (Source: MiniMax); we grade it T4 because its methodology is contested and not peer-reviewed.
The reporting bias. Published copy-trading track records are typically gross of fees, uncorrected for multiple testing across the platform's entire provider population, and subject to survivorship — providers who blow up are delisted, so the visible distribution is the surviving right tail of a much wider one (Source: MiniMax). A platform hosting 10,000 signal providers will display several with extraordinary records for the same reason that 10,000 coin-flippers produce several long runs of heads.
At USD 100 over 90 days. MiniMax puts P(USD 100 → USD 200 via a paid signal service) below 5% [T6] (Source: MiniMax). The figure is unmodeled. The stronger statement is structural and does not require a probability: a service advertising a USD 100 → USD 200 outcome in 90 days is, with high prior probability, monetizing the subscriber rather than the alpha, and the subscription fee itself is a certain negative return applied to a stake of USD 100. A USD 20 monthly subscription consumes 60% of the stake over the window before a single trade. Non-viable.
Verdict: ruin in finite time under a finite bankroll, with the ruin probability rising toward certainty in the trade count. The mathematics is not in dispute; both source reports mangled the worked example, so we rebuild it.
The mechanics. A martingale sizing rule doubles the stake after each loss, so the bet after k consecutive losses is b_{k+1} = 2^k · b₁ and the cumulative loss over those k bets is L_k = (2^k − 1) · b₁. The scheme's appeal is that a single win recovers the entire accumulated loss plus b₁. Its defect is that the required stake grows exponentially against a bankroll that does not.
The USD 100 absorption proof. With B₀ = USD 100 and b₁ = USD 1, the largest k satisfying 2^(k+1) − 1 ≤ 100 is k = 5: five consecutive losses cost USD 31 cumulative, and the sixth bet of USD 32 is affordable. On a seventh consecutive loss the required bet is USD 64 while remaining capital is USD 37, and the scheme fails outright [T1, derived] (Source: Gemini). Gemini states the streak count three different ways across its own deliverable — six, five-survivable-with-bankruptcy-on-the-seventh, and seven — and we resolve it to the derivation above: the account survives six bets and cannot fund the seventh.
The ruin probability. For a fair binary bet at p = 0.5 with the sizing above, the probability of encountering a ruinous seven-loss streak over N trades is approximately 1 − exp(−N · 0.5 · 0.5⁷). That gives P(ruin) ≈ 54.2% over N = 200 trades and ≈ 85.8% over N = 500, converging to 1.00 as N grows [T1, derived; arithmetic verified in digest] (Source: Gemini). Note the input: this is a fair game. The scheme does not require an unfavorable edge to destroy the account; it requires only enough repetitions. Introduce realistic transaction costs and p falls below 0.5, accelerating every figure above.
We discard MiniMax's competing worked example entirely. Its digest establishes that the example both mis-adds (1 − 0.49⁷ − 0.51⁷ = 98.4%, not the stated 98.7%) and describes the wrong mechanic — "sequential bets that each shrink the bankroll by half on a loss" is the inverse of martingale — so it does not illustrate the proposition it is attached to (Source: MiniMax — dropped).
Why the theory says do not play at all. Kelly (1956) establishes that optimal sizing under a positive edge is proportional to edge divided by variance [T1] (Bell System Technical Journal 35(4), 917–926) (Source: Gemini, MiniMax). For a game with no edge or a negative edge, the Kelly fraction is zero or negative — the optimal bet size is nothing. Doubling after a loss is not merely suboptimal under the log-growth criterion; it is the maximal-deviation-from-optimal response to a game one should not be playing. Dubins and Savage (1965) supply the complementary result: in a subfair game, bold play — wagering min(current wealth, amount needed to reach target) — strictly maximizes the probability of reaching a fixed target [T1] (How to Gamble If You Must: Inequalities for Stochastic Processes) (Source: Gemini, Qwen, MiniMax). Martingale is neither bold nor Kelly. It is a timid-play schedule with an exploding tail, which is the worst available combination for a fixed-target problem: it maximizes the number of trials — and therefore the cumulative ruin hazard computed above — while never concentrating enough stake into any single trial to move the target-hitting probability.
Anti-martingale and progressive schemes. Increasing size after wins rather than losses is not ruinous in the same finite-time sense, and Qwen states its actual defect correctly: it maximizes exposure at the point of maximum accumulated gain, so the inevitable end of a winning streak occurs against the largest position and surrenders a disproportionate share of the run (Source: Qwen). The general principle covering all progressive schemes is the one MiniMax states plainly and correctly: position sizing can amplify a positive edge but cannot manufacture one from a negative expectation (Source: MiniMax). Every sizing rule is a linear operator on the per-trade expectation; none changes its sign.
Why the scheme persists. Its survival in retail trading documents a behavioral regularity — the belief that an adverse run must reverse — rather than any property of the scheme. Qwen adds the practical constraint that even a fair game defeats martingale because players have finite wealth and venues impose position limits (Source: Qwen), which is the USD 100 absorption proof stated in words.
At USD 100 over 90 days. Structurally non-viable. The exposed quantity is the entire stake, the ruin probability exceeds 50% within 200 trades even at fair odds, and the scheme is simultaneously worse than doing nothing (negative EV amplified by trade count) and worse than bold play (fails to concentrate stake toward the target).
Verdict: the meta-cause of false confidence in every preceding subsection. The mechanism is not sloppiness — it is that the search procedure that finds a strategy is also the procedure that inflates its apparent performance, and the inflation is invisible from inside the search.
This subsection describes the generating mechanism. The remedies — purged and combinatorial cross-validation, deflated Sharpe, PBO, pre-registration — belong to Section 8 and are not restated here.
The central result. Bailey, Borwein, López de Prado and Zhu (2014) give the Minimum Backtest Length required to prevent a false discovery at a given trial count [T1] ("Pseudo-Mathematics and Financial Analytics," Notices of the AMS 61(5), 458–471) (Source: Gemini):
MBL > (2 ln N) / E[SR]² · (1 − γ₁·E[SR] + ((γ₂ − 1)/4)·E[SR]²)
Its practical implication is the single most important fact in this section. Testing N = 100 strategy variations on three years of daily data mathematically guarantees an in-sample Sharpe above 2.0 by chance alone, and preventing that false discovery at N = 100 requires more than twelve years of data (Source: Gemini). Gemini's own JSON weakens this to N ≥ 20 variations; we use the body figure of N = 100 and record the internal discrepancy. The point survives either number: the trial counts at which backtests become uninformative are trial counts a single person reaches casually, and nothing in the backtest output signals that the threshold was crossed.
The mechanisms, and what each inflates. MiniMax supplies the fullest taxonomy in the corpus. Its section is headed "Eight mechanisms" and enumerates ten; we present the ten and note the defect (Source: MiniMax, with Gemini's parallel ten-trap list).
Why this generates confidence rather than doubt. The inflation is not perceptible from inside the process. Each individual decision — trying a second lookback window, dropping a delisted name because its data is messy, using the vendor's adjusted price series — is locally reasonable, and none of them announces itself as a trial. The researcher's subjective count of hypotheses tested is systematically far below the true N, so even a researcher who intends to apply a multiple-testing correction applies it at the wrong N. The result is a backtest whose apparent quality rises monotonically with effort, which is precisely the feedback signal a diligent person will pursue.
Qwen's contribution. Qwen identifies six of these traps with mitigations and adds a useful reframing: the purpose of a 90-day live experiment is not to prove a strategy works — the statistical power is far too low for that — but to learn the practical challenges of execution and observe the real distribution of outcomes (Source: Qwen). Qwen omits deflated Sharpe, PBO and the White/Hansen family entirely, which is a gap against the other two reports.
At USD 100 over 90 days. MiniMax's summary judgment is the right one to carry forward: no strategy presented without purged cross-validation, walk-forward validation and a deflated-Sharpe adjustment should be treated as viable, and in-sample Sharpe should be assumed to overstate out-of-sample Sharpe by a factor of 2× to 5× at the typical retail backtest [T6 for the factor] (Source: MiniMax).
| # | Category | Primary failure mode | Best evidence | Tier | Direct USD 100 / 90-day cost |
|---|---|---|---|---|---|
| 7.1 | Technical analysis patterns | No signal post-1988; competed away | Park & Irwin (2007); Bajgrowicz & Scaillet (2012) | [T1] | Zero rules survive after 5–10 bps costs |
| 7.2 | Retail day trading | No edge; friction × turnover | Barber et al. (2014); Chague et al. (2020) | [T1] | 97% lose; <1% show persistent skill |
| 7.3 | Leveraged/inverse ETFs > 1 day | Deterministic path-dependent drag | Gemini derivation; Avellaneda & Zhang (2010) | [T1] | ≈ −4.6% over 90d at L=3, σ=25%, flat index |
| 7.4 | Penny stocks / OTC / pink sheets | Structural transfer via spread + manipulation | Bradley et al. (2014) | [T1] | 10–50% round-trip spread |
| 7.5 | Social-media / meme / sentiment | Signal lags price; unbounded feature space | Nofsinger et al. (2021); Pedersen (2022) | [T2] | One event window, no replication |
| 7.6 | Naive ML without purged CV | Label leakage; trial-count inflation | Hou et al. (2020); McLean & Pontiff (2016) | [T1] | IS Sharpe > 4.0 → ~0 live |
| 7.7 | Copy-trading / signal services | Equilibrium; latency; reporting bias | Carhart (1997); Fama & French (2010) by analogy | [T1] (analogy) / [T4] (direct) | Subscription alone can exceed 50% of stake |
| 7.8 | Martingale / progressive sizing | Finite-bankroll absorption | Gemini derivation; Kelly (1956); Dubins & Savage (1965) | [T1] | Ruin at 7th consecutive loss; P(ruin) 54% at N=200 |
| 7.9 | Overfit backtests | Search inflates its own result invisibly | Bailey et al. (2014); Harvey et al. (2016) | [T1] | N=100 trials on 3y data guarantees Sharpe > 2.0 by chance |
The unified finding. None of the nine categories carries positive expected value at USD 100 scale over 90 days under multiple-testing-corrected methodology. All three reports reach this conclusion independently, and none of the disagreements among them touches the sign of any verdict — they disagree on magnitudes, citations and mechanism details, never on direction (Source: Qwen, Gemini, MiniMax). MiniMax quantifies the ensemble as P(any one of the nine achieving P(USD 100 → USD 200) > 0.05) below 5% [T6] (Source: MiniMax); that is an unmodeled second-order estimate and we carry it as such.
Two asymmetries worth stating explicitly. First, our uncertainty about these verdicts is not symmetric. The probability that we are wrong about any individual category being negative-EV is low and roughly equal across categories; the probability that we are wrong about the magnitude of a specific cited figure is substantially higher, given how many source numbers failed verification. A reader should trust the signs far more than the decimals. Second, the categories fail for different reasons, which means they do not diversify against one another. Combining a technical-pattern entry rule with a martingale sizing rule and a sentiment filter does not produce three partial edges; it produces a strategy with no edge, no risk control, and an inflated backtest, which is strictly worse than any component alone.
Leveraged-ETF drag formula. MiniMax: drift penalty of −2σ² per day for a 2× fund (yielding 2.9% over 90 days at σ=20%, 6.5% at σ=30%). Gemini: X_t = X_0(S_t/S_0)^L · exp(½(L−L²)σ²t), the standard beta-slippage result, giving −4.6% over 90 days at L=3, σ=25% on a flat index. Resolution: Gemini. MiniMax's own digest identifies its formula as roughly 2× too large and contradicting the standard result; Gemini's arithmetic was independently verified in digest. Every MiniMax magnitude in that section is dropped.
Leveraged-ETF effect on P(reaching a fixed target). MiniMax: P(USD 100 → USD 200) via a 2× ETF is "essentially the same as 1× less the drag." Resolution: rejected as analytically wrong, not merely imprecise. A 2× fund needs +41% on the underlying, not +100%; leverage raises threshold-hitting probability while lowering expected value. The section now states this explicitly as the reason the instrument is seductive, which is a stronger and more accurate treatment than any single source provided.
Trainor (2010) leveraged-ETF underperformance of 25–75% per year. MiniMax (also referenced by Gemini). Resolution: dropped. Irreconcilable by an order of magnitude with the ~2–8% annual drag the shared formula implies for broad-index funds; MiniMax's own digest flags the internal contradiction and Gemini's digest could not confirm the publication venue.
Bajgrowicz & Scaillet rule universe. Gemini: 15,000+ rules, DOI 10.1016/j.jfineco.2012.06.002. MiniMax: 5,580 rules, DOI 10.1016/j.jfineco.2012.08.002, plus an unreconciled adjacent reference to Sullivan-Timmermann-White's 7,846. Resolution: irreconcilable on the count; both excluded. The finding both agree on — zero rules survive after 5–10 bps costs — is reported qualitatively. Gemini's DOI is used as the better-formed of the two.
Technical analysis: does anything survive multiple-testing correction? MiniMax asserts nothing survives, then four lines later reports that Sullivan-Timmermann-White's Reality Check results do survive. Resolution: unified temporally, not by majority. Rules survived correction on pre-1988 data (Brock-Lakonishok-LeBaron, confirmed by Sullivan-Timmermann-White); the effect vanished thereafter and fails after costs in modern data (Park & Irwin, Bajgrowicz & Scaillet). Gemini's institutional-algorithmic-arbitrage mechanism supplies the reconciling explanation. Both source claims are true of different samples.
Park & Irwin (2007) title and venue. MiniMax: "A Reality Check for Technical Trading Rules," venue hedged across three journals, RFS DOI. Gemini: "What Do We Know About the Profitability of Technical Analysis?", Journal of Economic Surveys 21(4), 786–826, DOI 10.1111/j.1467-6419.2007.00519.x. Resolution: Gemini. MiniMax's title collides with White (2000) and its own source entry hedges across three venues, which a real citation does not need to do.
Day-trader base rate — Taiwan. MiniMax: top 0.1% profitable, 99.9% lose, mean −0.26%/day; also states "approximately 1%" eight lines later. Gemini: 99% net unprofitable, <1% with repeatable skill. Resolution: neither figure as stated. Both digests establish that the paper's finding concerns persistent, predictable profitability, not the fraction losing money in a period. Reported as: fewer than 1% show persistent profitability net of fees. MiniMax's −0.26%/day and its internal 0.1%-vs-1% discrepancy are both excluded.
Day-trader base rate — Brazil (Chague et al.). MiniMax: 1.1% beat the CDI gross, 0.4% net. Gemini: 19,642 traders over 300 days, 97% lost money, 1.1% above minimum wage (~USD 54/day), 0.1% above USD 300/day. Resolution: Gemini. MiniMax's own digest flags its version as a re-benchmarking onto a comparator the study did not use. Note: the two reports also supply incompatible DOIs (Social Science Research prefix vs Brazilian Review of Finance); neither DOI is carried forward.
Barber & Odean (2000) characterization. MiniMax: "the median household lost money after costs," with 6.5 percentage points applied to the median active trader. Resolution: corrected. The paper reports underperformance (≈16.4% vs 17.9%) over a bull-market sample, and the 6.5-point gap belongs to the highest-turnover quintile. Stated correctly in 7.2. Gemini's bibliography supplied the correct citation (DOI 10.1111/0022-1082.00223) despite never citing it in its own body.
Settlement regime. Qwen: T+2, asserted three times and load-bearing for its turnover analysis. Gemini: T+1 under 17 CFR § 240.15c6-1. Resolution: Gemini. U.S. equities moved to T+1 in May 2024; Qwen's figure is stale for a 2026 report and materially understates achievable cash-account turnover.
McLean & Pontiff (2016) figures. MiniMax: in-sample abnormal returns of 26%/year decaying to 12%/year then 4%/year, an "85% decay." Gemini: 58% post-publication decay, decomposed as 26% overfitting + 32% arbitrage. Resolution: Gemini. MiniMax converted decay percentages into annual return levels, a category error; a 26%/year in-sample abnormal return across the anomaly universe is implausible on its face.
Hou, Xue & Zhang (2020) replication rates. MiniMax: 23% remained significant, 12% survived a stricter battery. Gemini: 65%+ fail at t ≥ 1.96, 82%+ fail at t ≥ 2.78. Resolution: Gemini, whose figures match the paper's standard reported results; MiniMax's own digest flags its version as not matching. Note the two also supply different DOIs (hhy099 vs hhy131); Gemini's is used.
Harvey, Liu & Zhu (2016) threshold framing. MiniMax quotes the finding as "t-statistics greater than 3.0 are not adequate," which contradicts its own following sentence. Gemini: t > 2.0 is inadequate; the corrected threshold is t > 3.0 across 315 factors. Resolution: Gemini. MiniMax's rendering is an inverted quotation.
Martingale worked example. MiniMax: 49% win rate, 1:1 binary, "probability of doubling 7 sequential bets," with arithmetic totalling 98.7%. Gemini: b_{k+1} = 2^k·b₁, k_max derivation on a USD 100 bankroll, P(ruin) ≈ 54.2% at N=200 and 85.8% at N=500. Resolution: Gemini's derivation adopted; MiniMax's example dropped entirely. MiniMax's arithmetic is wrong (correct value 98.4%) and its verbal description — bets that halve the bankroll on a loss — describes the inverse of martingale.
Martingale loss-streak count. Gemini states it three ways across its own deliverable: ruin within k=6, k_max=5 survivable with bankruptcy on the 7th, and "7 consecutive losses = complete bankruptcy." Resolution: stated once, from the derivation. With B₀=100 and b₁=1 the account funds six bets (cumulative USD 63) and cannot fund the seventh (USD 64 required against USD 37 remaining).
Penny-stock spread magnitude. MiniMax: 7–12% markup on pink sheets (Li & Zheng 2020). Gemini: 10–50% of share price (Bradley et al. 2014). Resolution: Gemini. MiniMax's source could not be placed by its own digest and is flagged as possibly fabricated; Gemini's range subsumes MiniMax's anyway.
Overfit-backtest trial threshold. Gemini's body: N = 100 variations on 3 years of daily data guarantees a chance Sharpe > 2.0, requiring 12+ years of MBL. Gemini's JSON: N ≥ 20. Resolution: the body figure (N = 100), with the internal discrepancy recorded. The argument is unaffected by which is used.
Coverage conflict — penny stocks. Qwen omits penny stocks, OTC and pink sheets entirely; Gemini and MiniMax both cover the category. Resolution: covered from Gemini and MiniMax; Qwen's omission recorded as a gap rather than as dissent.
Copy-trading citation base. All three reports assert the verdict; none supplies a verifiable direct citation. Gemini's JSON cites a topic label ("Empirical Market Microstructure Analysis"); MiniMax cites a working paper and three enforcement actions its own digest flags as unverifiable or fabricated; Qwen cites nothing. Resolution: verdict retained, all direct citations dropped. Rebuilt on the equilibrium argument, Gemini's microstructure/latency argument, and the fund-persistence literature (Carhart 1997; Fama & French 2010) as an explicitly labeled analogy, with the absence of direct evidence stated in the text.
"eRank" as a copy-trading platform. MiniMax lists it alongside ZuluTrade, NAGA and MT5. Resolution: excluded. eRank is an Etsy marketplace-analytics product, not a trading platform.
Citations downgraded or dropped as fabricated/unverifiable: Li & Zheng (2020) pink-sheet markup; Costanzino & Savona (2024) meme-stock abnormal returns and the −8.6%/36-month figure; Theouchi (2011) "Gambling Dynamically with a Variable Stage Cost"; Isiaq & Garcia-Bonete (2020) MQL5 signals; "SEC v. Managed Funds Association," "SEC v. TODAQ (2023)," and the CFTC "CZBCrypto" case; the "Securities Enforcement Commission" and its 31%-of-fraud and 270-issuer figures; the SEC September 2020 COVID retail-trading staff report and its ~9% aggregate loss figure; Ince & Porter's "Score of the SIR (Survivorship Index Revision)"; MiniMax's √(k/T) overfitting inflation formula; Gemini's JSON citations "Empirical Market Microstructure Analysis" and "Probability Theory & Markov Chains" (the latter carrying a spurious T1 grade); Trainor (2010)'s 25–75% figure; Da/Engelberg/Gao's 0.22%-per-SD point estimate (downgraded to T6, paper retained at T1); Brock-Lakonishok-LeBaron's "approximately 1,200 citations."
Section word count: 8,631 words — measured, body through §7.10 inclusive of the comparative table, excluding this Resolved Conflicts subsection. Whole-file count including the Resolved Conflicts audit trail: 10,111 words.
The length-parity requirement against Section 5 is not yet verifiable from here: Section 5 had not been written when this section was completed, and a request to its author for a target count went unanswered. Both figures above are reported so the comparison can be made downstream. Note for whoever performs it: the 8,631 body figure is the like-for-like comparator, since Section 5 is not expected to carry an equivalent conflict ledger.
Evidence-tier tag counts (body prose, bracketed tags): [T1] 40 · [T2] 15 · [T3] 2 · [T4] 7 · [T5] 5 · [T6] 16.
A backtest is not evidence. It is a hypothesis generator whose output distribution, under the null of zero skill, is systematically positive. Every methodological trap catalogued below shares one mechanism: it inflates the apparent in-sample performance of a rule that has no out-of-sample edge, and it does so in a direction the researcher wants to believe. All three source reports converge on this framing without contradiction (Source: Qwen, Gemini, MiniMax).
The framing that matters for a USD 100 → USD 200 / 90-day objective
is narrower than the general literature. Under a deadline, the quantity
being estimated is a first-passage probability, and first-passage
probabilities are nonlinear in the signal-to-noise ratio μ/σ. MiniMax
makes the consequence explicit: the overstatement is convex, so the
worse the true signal-to-noise, the larger the relative
inflation of P(reach target). A backtest reporting P(reach USD 200) =
0.30 may correspond, after deflation for selection, to a true
probability below 0.05 [T6] — the report offers no
derivation for this claim, so it is carried here as a directional
intuition, not a result (Source: MiniMax).
The engine underneath every trap in this section is order statistics. The expected maximum of N i.i.d. standard normal variates is approximately
E[max_N] ≈ √(2 ln N) − (ln(ln N) + ln(4π)) / (2 √(2 ln N))
[T1] (Lo, A. W., 2002, "The Statistics of Sharpe
Ratios," Financial Analysts Journal 58(4),
doi:10.2469/faj.v58.n4.2453) (Source: MiniMax). For N = 1,000
candidate rules, the uncorrected leading term √(2 ln 1000) = 3.72 and
the corrected expression evaluates to 3.12
[T3]. MiniMax reports 3.26 for this quantity; that figure
reproduces from neither the leading term nor the corrected formula, and
is not carried forward. More importantly, this expression returns the
expected maximum of N standard normals, not of N Sharpe
ratios. Converting requires multiplying by the standard error of
the Sharpe estimator (≈ 1/√T), so any statement of the form "the
expected maximum Sharpe is ≈ 3.2" is dimensionally incomplete without a
stated sample length [T3] (Source: MiniMax,
corrected).
The practical rule of thumb, stated by MiniMax and consistent with
the same order statistics, is E[max Z | N] ≈ √(2 ln N)
[T3] (Source: MiniMax). Search 100 variants and
you should expect a best-of-sample z-statistic near 3.0 from noise
alone.
Gemini supplies the complementary inversion — the Minimum Backtest Length (MBL) required before a reported Sharpe survives the search:
MBL > (2 ln N) / E[SR]² · ( 1 − γ₁·E[SR] + ((γ₂ − 1)/4)·E[SR]² )
[T1] (Bailey, D. H., Borwein, J., López de Prado, M.,
& Zhu, Q. J., 2014, "Pseudo-Mathematics and Financial Charlatanism:
The Effects of Backtest Overfitting on Out-of-Sample Performance,"
Notices of the AMS 61(5), 458–471, doi:10.1090/noti1105)
(Source: Gemini, MiniMax). Gemini's operational reading:
testing N = 100 variations on three years of daily data
mathematically guarantees an overfit Sharpe above 2.0 by
chance, and roughly 12+ years of data are
required to prevent false discovery at that search intensity
[T1] (Source: Gemini). Gemini's own JSON appendix
states a materially weaker threshold ("N ≥ 20 variations"); the two
numbers are not reconciled inside the source, and the N = 100 / 12-year
form is the one carried here as the better-specified of the two.
The asymmetry a reader should hold: this mechanism produces false positives, essentially never false negatives. A search that finds nothing is informative. A search that finds a 3-Sharpe rule is not.
Ten named traps. MiniMax addresses all ten with detection method, mitigation, and tooling; Gemini addresses all ten in prose (five in its machine-readable appendix); Qwen addresses six (Source: Qwen, Gemini, MiniMax). Where the three disagree on mitigation specifics, the resolution is recorded in §8.6.
A.1 Lookahead bias — using information not available
at the decision timestamp. MiniMax names four retail manifestations:
retroactively split-adjusted close prices applied to pre-split decision
timestamps; "most-recent" fundamental values that have since been
restated; index membership rebalanced after the backtest window; and
earnings surprises computed against a consensus not yet aggregated at
the decision timestamp [T3] (Source: MiniMax).
Detection: interrogate every input — does the dataset at
timestamp t contain only information published by t?
Run the backtest twice, once with strictly lagged inputs and once with
naïvely aligned inputs, and compare [T3] (Source:
MiniMax). Gemini's version is more mechanical: check whether the
feature matrix X_t contains bar-t close/high/low
or unannounced fundamental filings [T3] (Source:
Gemini).
Mitigation: shift the feature matrix by at least one lag
(X_{t−1} → R_t) and index SEC data by filing
acceptance timestamp t_file, not period end
[T3] (Source: Gemini). MiniMax's lag conventions:
one trading day for prices, one business day for fundamentals, one
quarter for fundamental filings to absorb the SEC reporting lag
[T5] (Source: MiniMax). Note that MiniMax states
this lag as a flat "45 days"; SEC 10-Q deadlines are 40 or 45 days
depending on filer status, so treat 45 days as the conservative bound,
not the rule.
Python: pandas.merge_asof with strict
inequality joins; polars.shift(1); pydantic
for schema-level enforcement of as-of columns; pyarrow
parquet with explicit as-of columns (Source: Gemini,
MiniMax).
A.2 Survivorship bias — applying a current-universe ticker list to historical data, silently dropping delisted, acquired, renamed, and bankrupt names.
Detection: re-run on a delisting-aware price file and
compare equity curves; report the ratio of surviving to delisted names
by year [T3]; compare the active constituent list against
historical delisting archives as of date t [T3]
(Source: Gemini, MiniMax). MiniMax supplies the single most
usable threshold in this section: if the Sharpe changes by more
than 0.3 when a delisting-return adjustment is applied, the backtest was
substantially contaminated [T3] (Source:
MiniMax).
Mitigation: a delisting-adjusted dataset — CRSP/Compustat or
Norgate point-in-time constituent archives including delisting returns
R_delist [T4] (Source: Gemini,
MiniMax). The free substitute is a manual construction from SEC
EDGAR filing headers carrying delist_date and
effective_date, with every row tagged
asof_date [T3] (Source: MiniMax).
The foundational citation is Brown, S. J., Goetzmann, W., Ibbotson,
R. G., & Ross, S. A. (1992), "Survivorship Bias in Performance
Studies," Review of Financial Studies 5(4), 553–580,
doi:10.1093/rfs/5.4.553 [T1] (Source: MiniMax).
MiniMax attaches a magnitude to this paper — that survivor-only
mutual-fund samples inflate average returns by 1.0–1.5% per
annum — but BGIR (1992) is principally a theoretical and
simulation treatment of survivorship-induced spurious persistence, and
percentage-per-annum magnitudes of this order are conventionally sourced
elsewhere. The magnitude is carried here as [T6]
with the attribution explicitly in doubt, and should not be
quoted against BGIR (Source: MiniMax, downgraded).
At a USD 100 stake the absolute dollar consequence of survivorship
contamination is trivial. The consequence for the inference is
not: MiniMax's point is that the distortion is large enough to flip a
deflated Sharpe from positive to negative [T3] (Source:
MiniMax).
A.3 Selection bias — choosing a backtest period, asset universe, or parameter range after observing which slice produces a positive result.
Detection: there is no post-hoc detection. The only
instrument is pre-commitment: pre-register period, universe, and
parameter grid before running, and report results on a held-out slice
[T3] (Source: Gemini, MiniMax, Qwen). White's
reality check (§B.6) is the retrospective correction for a candidate
pool of known size, and it requires knowing N honestly.
Mitigation: a strict out-of-sample window never re-used once
a result has been observed, plus an appendix documenting all rejected
parameter sets [T3]. MiniMax's retail-practical substitute
for academic pre-registration: maintain a decisions.log
recording strategy and parameters before the backtest runs,
with an advance commitment to report every pre-registered strategy
including the failures [T3] (Source: MiniMax).
This is the cheapest high-value control in the entire section and the
one most reliably skipped.
Qwen omits this trap's detection machinery beyond naming pre-registration and multiple-testing correction (Source: Qwen).
A.4 Data-snooping and multiple testing — testing many rules and reporting the best as though it were the only one tested.
Detection: report the number of independent trials N;
compute the deflated Sharpe ratio; run White's reality check or Hansen's
SPA against the candidate pool [T3]. Gemini's operational
form: calculate the trial count N and evaluate the variance of
the Sharpe ratios across trials, V[{SR_k}] — which is the
quantity the DSR actually needs [T3] (Source: Gemini,
MiniMax).
Mitigation: family-wise error rate or false-discovery-rate
correction on the trial pool — Bonferroni (most conservative), Holm
(step-down, less conservative), Benjamini–Hochberg (controls FDR),
Romano–Wolf (bootstrap-based, controls FWER under cross-strategy
dependence) [T3]. Report p-values adjusted for N
candidates, never raw (Source: Gemini, MiniMax, Qwen — Qwen names
only Benjamini–Hochberg). MiniMax attaches two DOIs to this menu
(10.1214/aos/1176348899, 10.1093/rfs/4.4.867)
that appear nowhere in its own bibliography and do not match the methods
named; both are dropped as unverifiable and the
correction menu is carried on the strength of the method names
alone.
The headline empirical demonstration is Sullivan, R., Timmermann, A.,
& White, H. (1999), "Data-Snooping, Technical Trading Rule
Performance, and the Bootstrap," Journal of Finance 54(5),
1647–1691, doi:10.1111/0022-1082.00163 [T1] — 7,846
simple technical trading rules on 100 years of Dow data; after
correcting for the search, the best rules' apparent t-statistics
collapsed to insignificance (Source: MiniMax). Gemini
contributes the modern replication: Bajgrowicz, P., & Scaillet, O.
(2012), Journal of Financial Economics 106(3), 473–491,
doi:10.1016/j.jfineco.2012.06.002 [T1] — under White's
Reality Check plus FDR correction and 5–10 bps of costs, zero
rules generate significant out-of-sample excess returns
(Source: Gemini).
MiniMax's redefinition of N is the most decision-relevant claim in
this subsection and is carried forward verbatim in substance:
for a 90-day experiment, the relevant N is not the number of
rules but the number of independent decisions — at most ~63 for
a daily-rebalanced strategy, the count of distinct contracts for an
event-contract strategy, the count of entry signals for an options
strategy [T3] (Source: MiniMax). This is what
makes §8.5's power arithmetic binding rather than academic.
A.5 Overfitting (parameter and feature) — fit flexibility exceeding the information content of the sample.
Detection: purged and embargoed k-fold cross-validation;
report in-sample Sharpe, out-of-sample Sharpe, and the ratio; evaluate
the IS-vs-OOS Sharpe gap and the combinatorial-CV rank [T3]
(Source: Gemini, MiniMax).
Mitigation: reduce degrees of freedom and penalize the
search. Gemini: L1/L2 regularization and tree depth ≤ 3
for tree-based learners, plus a PBO calculation [T3]
(Source: Gemini). MiniMax: limit the count of engineered
features to O(√T), where T is the number of
independent returns [T6] — uncited, and stated
inconsistently within MiniMax itself, which elsewhere applies the same
√T bound to the count of candidate strategies rather than
features. Both forms are carried, both as [T6], because the
bound is a heuristic with no source in any of the three reports
(Source: MiniMax, downgraded).
MiniMax attributes a simulation result to Bailey et al. (2014) — that
1,000 candidate rules tested on a 4-year window with true Sharpe 0
produce a selected-rule OOS Sharpe averaging +0.6 with
standard deviation ~0.6, run on a simulator it names "MinnSim." Neither
the simulator name nor the figures could be corroborated against the
cited paper. Carried as [T6], flagged, and not to
be quoted as a peer-reviewed result (Source: MiniMax,
downgraded).
Qwen's treatment of overfitting reaches only "out-of-sample testing,
regularization, triple-barrier labeling," with no quantified stopping
rule [T4] (Source: Qwen).
A.6 Regime change and non-stationarity — parameters
calibrated on a regime that no longer obtains. Qwen states the mechanism
most plainly: financial time series are non-stationary, so mean and
variance change over time, and naive machine learning on price series
fails for this reason before any other [T4] (Source:
Qwen).
Detection: test parameter stability across rolling windows;
apply Chow or Quandt–Andrews breakpoint tests; compute CAGR, Sharpe, and
tail risk per regime [T3] (Source:
MiniMax).
Mitigation: three distinct proposals, all compatible and all
retained. Gemini: Hidden Markov Model regime gating
plus crisis sub-sample stress testing [T3] (Source:
Gemini). MiniMax: walk-forward with re-estimation frequency matched
to the natural regime length, ensembling across regimes, and a hard
acceptance criterion — require a consistent sign of edge in at
least two of three non-overlapping periods [T3]
(Source: MiniMax). Qwen: walk-forward analysis, regime-aware
models, stress testing across environments [T4]
(Source: Qwen).
Citation: Ang, A., & Bekaert, G. (2002), "International Asset
Allocation With Regime Shifts," Review of Financial Studies
15(4), 1137–1187, doi:10.1093/rfs/15.4.1137 [T1]
(Source: MiniMax).
Python: ruptures for change-point detection
(CUSUM, PELT, BinSeg) applied both to the strategy P&L series and to
each input feature; statsmodels for Chow and Andrews tests;
arch for GARCH-derived regime indicators (Source:
MiniMax).
A.7 Transaction-cost underestimation — assuming zero
commissions, zero slippage, zero market impact. At a USD 100 stake this
is the trap with the largest dollar consequence, because cost
is dominated by spread and fixed fees, and the spread is a far larger
fraction of a USD 100 trade than of a USD 1M trade [T3]
(Source: MiniMax).
Detection: re-run with explicit per-trade cost — commissions
plus exchange fees plus half-spread plus temporary impact — and compute
round-trip cost as a percentage of stake; verify the gross-to-net Sharpe
decay [T3]. Gemini: compare mid-price execution against the
full bid-ask spread and the exchange taker-fee schedule
[T3] (Source: Gemini, MiniMax).
Mitigation: Gemini's decomposition
Cost = Spread/2 + Slippage + Fees [T3]
(Source: Gemini); Qwen's is the same idea less formally
("incorporate realistic commissions, fees, and bid-ask spreads")
[T4] (Source: Qwen). MiniMax adds two acceptance
rules worth adopting because they are falsifiable: require net
Sharpe > 0 under conservative costs, and report sensitivity to a 2×
cost assumption — a strategy whose Sharpe goes negative at
twice the assumed cost is not robust [T3]. And: a
backtest that does not report per-trade cost in basis points, with the
cost model stated, is unverified [T3] (Source:
MiniMax).
Citation dropped. MiniMax cites Frazzini, Israel & Moskowitz, "Trading Costs of Asset Pricing Anomalies" (doi:10.2139/ssrn.2294498) for the claim that "13 of 17 well-known anomalies have a net-of-cost Sharpe of zero or negative." The paper, checked against its primary source, concludes the opposite: using nearly a trillion dollars of live trading data, real-world trading costs are "less than a tenth as large as previous studies suggest," strategy capacity is "more than an order of magnitude larger" than prior work indicated, and "the main anomalies to standard asset pricing models are robust, implementable, and sizeable." The claim is excluded from this report in every form. The transaction-cost trap stands on mechanism and on the retail-scale spread arithmetic, not on this citation (Source: MiniMax, excluded).
A.8 Liquidity and market-impact assumptions invalid at retail scale — assuming execution at historical VWAP when the order is a non-trivial fraction of average daily volume.
Detection: compute the median and 95th-percentile
participation rate against 20-day ADV; estimate impact from the
square-root law, κ·σ·√(Q/ADV) [T6] —
MiniMax presents this law with no citation and simultaneously instructs
the reader to "document the assumed κ" while supplying no κ value or
source, so the functional form is usable and the calibration is not
(Source: MiniMax).
Mitigation: cap participation at ≤ 1% of bar volume
/ ADV — Gemini and MiniMax agree on the 1% figure, with MiniMax
also floating a looser 1–5% band elsewhere; the 1% ceiling is the
resolved value [T3] (Source: Gemini, MiniMax).
Restrict to the most liquid ETF and equity subset; MiniMax's universe
floor is ADV > USD 1M for a 90-day experiment
[T3]. Gemini adds fractional-share routing penalties as an
explicit cost line [T3] (Source: Gemini).
MiniMax concedes honestly, and this concession should be preserved:
at USD 100 of capital, market impact is usually
negligible on highly liquid instruments (mega-cap US equities,
BTC, ETH, SPY, QQQ, TLT, GLD) and binds only on small-cap equities,
small-cap ETFs, altcoins, and thin-book event contracts
[T3] (Source: MiniMax). This is the one trap in
the list where the retail scale is a genuine advantage — and it is
exactly offset by A.7, where retail scale is a genuine disadvantage.
Qwen does not address this trap at all (Source: Qwen —
gap).
A.9 Point-in-time data failures and restatement contamination — using the most-recently-reported figure for a fundamental that has since been restated.
MiniMax's worked example: an analyst running a value strategy on 31
December 2008 uses the most recent reported book value per share, which
reflects impairments and write-downs not filed until the 2009 10-K. The
strategy trades on a "ghost" value [T3] (Source:
MiniMax).
Detection: compare every input row against the original SEC
filing date; for restated values, attach the original filing date as the
as-of and ignore later revisions inside the in-sample period
[T3] (Source: MiniMax).
Mitigation: use the earliest available
EDGAR filing of each value, not the latest. MiniMax's join rule, stated
precisely: maintain a vintage table keyed by filed_at, each
value carrying filed_at and value columns;
join the strategy to the value on
decision_date >= filed_at, taking the most recent
filed_at <= decision_date [T3] (Source:
MiniMax). Gemini's equivalent: raw as-reported EDGAR filings
indexed by acceptance timestamp [T3] (Source:
Gemini). Qwen names PIT data as the mitigation for survivorship
bias but does not treat restatement contamination as a distinct failure
[T4] (Source: Qwen — partial gap).
Canonical reference implementation: Croushore, D., & Stark, T.
(2001), "A real-time data set for macroeconomists," Journal of
Econometrics 105(1), 111–130, doi:10.1016/S0304-4076(01)00072-0
[T1] — the Philadelphia Fed Real-Time Data Set (Source:
MiniMax).
A.10 Backfill bias in vendor datasets — a vendor
retroactively adds new listings, splits, or index constituents to the
historical bar series stamped with original-event timestamps, when no
participant could have traded at that price at that time. MiniMax's
example: a vendor adds a stock in 2020 and backfills its price history
to 2010; the history looks complete in 2026, but a researcher who
downloaded the same feed in 2012 would never have seen those bars
[T3] (Source: MiniMax).
Detection: compare the current vendor universe against
contemporaneous vendor snapshots and identify bars absent from the
original release; audit vendor schema release histories
[T3] (Source: Gemini, MiniMax).
Mitigation: use a vendor that preserves vintage history —
FRED/ALFRED for macro, CRSP for equities — or maintain timestamped
static archives locally; for equities, build the historical universe
from EDGAR and refuse to add an asset before its first SEC filing date
[T3] (Source: Gemini, MiniMax). MiniMax's
distinctive observation, which no other source makes: this is
the only trap in the list that requires an internal archive rather than
a third-party product [T3]. You cannot buy your
way out of it retroactively. Qwen does not address this trap
(Source: Qwen — gap).
A.11 The data-availability constraint behind traps A.1, A.2,
A.9, and A.10. Four of the ten traps are only fixable with
point-in-time, survivorship-bias-free data, and MiniMax's assessment of
what is obtainable free is blunt: no major free retail endpoint
is both PIT and survivorship-bias-free for equities
[T3]. SEC EDGAR is PIT by construction for
fundamentals (every filing carries an explicit
filed_at; rate limit 10 requests/second with a descriptive
User-Agent header) [T5], and ALFRED and the Philadelphia
Fed Real-Time Data Set are PIT by construction for macro
[T5]. None of the three supplies delisting-aware
prices. MiniMax marks EDGAR "survivorship-bias-free: yes" and
then concedes two sections later that a delisting-aware price feed must
come from a paid vendor — an internal contradiction; the
resolved position is that EDGAR is PIT and survivorship-safe for filings
only, not for prices (Source: MiniMax, corrected).
Gemini independently flags yfinance as PIT = NO,
survivorship-free = NO, with terms-of-service restrictions on automation
[T5] (Source: Gemini), which MiniMax corroborates
and calls "the most common cause of survivorship bias in retail
backtests" [T3].
Paid PIT equity data runs USD 100–500/month minimum
— for a USD 100 experiment, the data required to make the backtest
honest costs more than the capital at risk, every month
[T4] (Source: MiniMax). Building the free
substitute (EDGAR-derived universe with first_trade_date
and delist_date, joined to a delisting-aware price series
and fundamentals by filing date) is, in MiniMax's words, "a non-trivial
engineering project that consumes the entire 90-day budget before any
backtest runs" [T3].
B.1 Purged k-fold cross-validation with embargo.
Standard scikit-learn k-fold assumes sample independence. Financial
labels span multiple observations — a triple-barrier label consumes the
next N bars — so a naive fold boundary leaks the test-fold outcome into
the training set [T3] (Source: Gemini, MiniMax,
Qwen).
Procedure: for each test fold t, retain a training
observation i if it lies entirely after the
embargoed test window or entirely before it — that is,
t_start[i] >= t_end[test_t] + embargo
OR
t_end[i] <= t_start[test_t] − embargo. Everything else
is purged. This corrects MiniMax's printed rule, which states
the condition as a conjunction (AND); as printed, no observation can
satisfy both branches and the entire training set is purged
[T3] (Source: MiniMax, corrected).
Citation: López de Prado, M. (2018), Advances in
Financial Machine Learning, Wiley, Chapter 7 [T4]
(Source: Gemini, MiniMax, Qwen). MiniMax offers
doi:10.1080/14697688.2019.1703030 (Quantitative Finance, 2020)
as "the most authoritative peer-reviewed reference"; that item is a
book review, and the same source list records López de
Prado as the author of a review of his own book. Downgraded —
purged k-fold rests on [T4] book provenance, not on a
peer-reviewed methodological validation, and this report will not claim
otherwise (Source: MiniMax, downgraded).
Python: Gemini points to a custom purged-CV splitter built
on scikit-learn's BaseCrossValidator, used
with nautilus_trader for execution [T4]
(Source: Gemini). MiniMax recommends
purgedcv >= 0.1.3 (PyPI, scikit-learn-compatible), but
its own text dates that release to the same day as the report, describes
it as a mature mlfinlab replacement, and elsewhere concedes
it does not expose the functions it is recommended for.
purgedcv is carried as [T6] — verify
independently before adopting. Both sources agree
mlfinlab is not installable from PyPI [T6]
(Source: Gemini, MiniMax).
B.2 Combinatorial purged cross-validation (CPCV).
Generalizes purged k-fold by generating all combinatorial
train/test splits. For N folds with k test folds per split, the number
of backtest paths is C(N, k); for N = 6, k = 2, that is
15 paths (verified: C(6,2) = 15)
[T1] (Bailey, Borwein, López de Prado & Zhu,
Journal of Computational Finance 20(2),
doi:10.21314/JCF.2016.322). Each path trains on the union of the
non-test folds and tests on each of the k test folds, yielding an
empirical distribution of out-of-sample performance rather than
a single point estimate — which is the entire point (Source: Gemini,
MiniMax).
Citation: Bailey, D. H., Borwein, J., López de Prado, M.,
& Zhu, Q. J., Journal of Computational Finance 20(2),
doi:10.21314/JCF.2016.322 [T1] (Source: MiniMax).
Gemini cites the same work as "Bailey et al. (2017), Journal of
Computational Finance 20(4), 39–70" with DOI
10.2139/ssrn.2326253 — an SSRN working-paper DOI, not the
journal's. Resolved to MiniMax's DOI, which carries the correct
JCF registrant prefix.
Terminological caution: MiniMax conflates CPCV
(combinatorial purged cross-validation) with CSCV
(combinatorially symmetric cross-validation) in at least one
place. These are distinct procedures; PBO in Bailey et al. is defined
via CSCV [T3]. Anyone implementing from a
description should confirm which one they are building (Source:
MiniMax, flagged).
B.3 Walk-forward analysis. Optimize on a rolling
training window, test on the immediately following window, roll both
forward. Out-of-sample performance is the concatenation of test-window
returns; the number of forward steps equals the number of test windows
[T4] (Source: MiniMax, Qwen).
Citation: Pardo, R. (2008), The Evaluation and
Optimization of Trading Strategies, 2nd ed., Wiley, ISBN
978-0-470-12801-5 [T4] — Gemini's bibliographic
form is the correct one and supersedes MiniMax's "Design, Testing, and
Optimization of Trading Systems," which is the 1992 first-edition title
attached to a 2008 date (Source: Gemini, MiniMax; resolved
to Gemini).
Named failure mode: re-optimizing too frequently inflates
in-sample fit at the expense of out-of-sample performance. Match
optimization frequency to the strategy's signal half-life
[T3] (Source: MiniMax).
Python: no canonical library. MiniMax names the pattern via
backtesting's Backtest.optimize with a custom
re-optimization schedule, or bt with a custom
bt.Algo [T6] (Source: MiniMax).
Gemini's stack points instead at vectorbt (vectorized) and
nautilus_trader (event-driven) as the backtest engines
[T4] (Source: Gemini).
Gemini's Cluster 6 prose does not name walk-forward analysis among its remedies — it appears only via the orphaned Pardo bibliography entry (Source: Gemini — gap).
B.4 Deflated Sharpe ratio (DSR). The single most
useful number a backtester can compute, and the one most often absent.
It adjusts an observed Sharpe for four separate inflation sources:
(i) the number of trials N, (ii) return non-normality via
skewness γ₃ and excess kurtosis γ₄, (iii) sample length T, and (iv)
return autocorrelation (Source: Gemini, MiniMax).
Gemini describes it as deflating the Probabilistic Sharpe Ratio
for N, γ₃, and γ₄ — consistent, and the more accurate framing of the
construction [T3] (Source: Gemini).
Citation: Bailey, D. H., & López de Prado, M. (2014),
"The Deflated Sharpe Ratio: Correcting for Selection Bias, Backtest
Overfitting, and Non-Normality," Journal of Portfolio
Management 40(5), 94–107, doi:10.3905/jpm.2014.40.5.094
[T1]; preprint SSRN 2460551, doi:10.2139/ssrn.2460551
[T3]. MiniMax's author order and full bibliographic
record are correct and supersede Gemini's "De Prado & Bailey
2014" (Source: MiniMax; resolved against Gemini).
Formula — corrected. MiniMax prints a DSR expression that is
not the Bailey–López de Prado deflated Sharpe ratio: it
applies the skew/kurtosis bracket to S₀ inside the numerator with no
counterpart in the source, adds a spurious −((2γ̂₄ − 3)/8)Ŝ²
term to the denominator, and then multiplies the whole expression by
1/√V[Ŝ], double-counting a variance normalization already
performed. It further defines the benchmark as
S₀ = √(V[max Ŝ]), which is wrong. The published form
is:
DSR = Z[ (Ŝ − S₀) · √(T − 1) / √( 1 − γ̂₃·Ŝ + ((γ̂₄ − 1)/4)·Ŝ² ) ]
S₀ = √( V[{Ŝₙ}] ) · ( (1 − γ)·Z⁻¹[1 − 1/N] + γ·Z⁻¹[1 − 1/(N·e)] )
where Z[·] is the standard normal CDF, γ is
the Euler–Mascheroni constant (≈ 0.5772), V[{Ŝₙ}] is the
variance across the N trials' Sharpe ratios (not the
variance of a maximum), T is the number of return observations, γ̂₃ is
sample skewness and γ̂₄ sample excess kurtosis. Output is a p-value
against the null of zero true Sharpe after correcting for the
search [T1] (Bailey & López de Prado, Journal of
Portfolio Management 40(5), 94–107, doi:10.3905/jpm.2014.40.5.094)
(Source: MiniMax as printed — corrected here; Gemini's V[{SR_k}]
detection note independently corroborates the S₀ construction).
Anyone porting MiniMax's printed formula will produce wrong
numbers; this matters because computing the DSR is the top operational
instruction in that report.
Python: no maintained first-party implementation as of
2026-08-01. Gemini's answer is scipy.stats plus a custom
DSR module [T4]; MiniMax's is a direct port from the 2014
appendix in numpy + scipy, estimated at
roughly 60 lines [T6]. The two agree that
it must be hand-rolled (Source: Gemini, MiniMax).
B.5 Probability of backtest overfitting (PBO). The
probability that the best in-sample strategy underperforms the
median out-of-sample strategy. Computed by running CSCV/CPCV
across the full candidate set, recording each configuration's in-sample
rank and out-of-sample rank, and counting the fraction of configurations
in which the in-sample best lands below the OOS median [T1]
(Bailey, Borwein, López de Prado & Zhu, Journal of Computational
Finance 20(2), doi:10.21314/JCF.2016.322) (Source: Gemini,
MiniMax).
Citation: doi:10.21314/JCF.2016.322 [T1];
supporting working paper Bailey et al., SSRN 2568435,
doi:10.2139/ssrn.2568435 [T3] (Source:
MiniMax).
Interpretation: PBO = 0.5 means in-sample selection
has no better than coin-flip out-of-sample value
[T3]. MiniMax additionally claims PBO → 1 as N/T → ∞ and
PBO > 0.5 for N/T > 0.5 "even with mild overfitting"; no
derivation is given and the 0.5 threshold carries more precision than
the cited appendix is shown to support — carried as
[T6] (Source: MiniMax, downgraded).
DSR vs PBO — the distinction worth internalizing: DSR tests
whether the best in-sample Sharpe differs significantly from zero after
correcting for the search. PBO tests whether the best in-sample strategy
is the same strategy that is best out-of-sample. A backtest can
pass DSR and fail PBO, in which case the strategy
family has average skill but the in-sample search has identified a
spuriously best rule [T3] (Source: MiniMax).
Report both; neither substitutes for the other.
Python: scipy.stats plus custom code, alongside
a purged-CV splitter [T4] (Source: Gemini,
MiniMax).
B.6 White's reality check and Hansen's SPA test.
White, H. (2000), "A Reality Check for Data Snooping,"
Econometrica 68(5), 1097–1126, doi:10.1111/1468-0262.00152
[T1] — tests the null that the best of N
candidate strategies has zero expected out-of-sample performance, by
bootstrapping the joint distribution of the candidates' performance
under the null. The bootstrap preserves the dependence structure across
candidates, which is essential because most candidate strategies are
highly correlated with one another [T1] (Source:
MiniMax; corroborated by Gemini, which cites White (2000) in its
bibliography and invokes the Reality Check by name in its
Bajgrowicz–Scaillet discussion).
Hansen, P. R. (2005), "A Test for Superior Predictive
Ability," Journal of Business & Economic Statistics 23(4),
365–380, doi:10.1198/073500105000000063 [T1] —
generalizes White by testing whether the best candidate significantly
beats a benchmark, rather than whether it significantly
beats zero. SPA is the more useful of the two when a
specific benchmark exists (for a USD 100 experiment, the benchmark is
buy-and-hold SPY, not zero) [T1] (Source: MiniMax; the
identical DOI appears in Gemini's bibliography, though Gemini never
invokes the test in its body).
Report both together. The inner resampling is the Politis–Romano
stationary bootstrap (§B.7) [T3].
Python — corrected. MiniMax states that SPA is "not in any
single first-party library" and that implementation "requires
numpy and scipy directly," hedging that
arch provides bootstrap utilities "in some versions."
This is wrong. The arch package has
provided first-class SPA, StepM, and
MCS classes in arch.bootstrap for years;
Hansen's SPA test is directly callable and should not be reimplemented
[T5] (Source: MiniMax, corrected). Gemini's stack
independently includes arch (Gemini pins 7.2.0; MiniMax
pins ≥ 8.0 — both version claims are unverified
[T6]; the package identity is what matters, not the
pin).
Neither White's Reality Check nor Hansen's SPA appears anywhere in Qwen's report, despite both being named requirements (Source: Qwen — gap).
B.7 The stationary bootstrap. Politis, D. N., &
Romano, J. P. (1994), "The Stationary Bootstrap," Journal of the
American Statistical Association 89(428), 1303–1313,
doi:10.1080/01621459.1994.10476870 [T1] (Source:
MiniMax; the same DOI appears as an uncited bibliography entry in
Gemini).
A block bootstrap with random block length, geometrically
distributed with mean q. Unlike the circular block bootstrap it
produces stationary replicates, making it the appropriate resampler for
time series with unknown dependence structure [T1].
Recommended tuning: q = O(n^{1/3}) for weakly dependent
series [T3].
Python: arch.bootstrap.StationaryBootstrap; the
package default block length is int(np.ceil(n**(1/3))),
matching the recommendation [T5] (Source: MiniMax —
independently confirmed).
Absent from Qwen entirely (Source: Qwen — gap).
B.8 Triple-barrier labeling with meta-labeling.
López de Prado, M. (2018), Advances in Financial Machine
Learning, Chapter 3 [T4] (Source: Gemini, MiniMax;
Qwen names triple-barrier labeling once, uncited).
Triple-barrier: label an observation at time t by
which of three barriers is hit first — an upper profit-take barrier, a
lower stop-loss barrier, or a vertical time-exit barrier — yielding a
categorical {−1, 0, +1} or binary ±1 label. The rationale is that this
represents the actual decision problem more faithfully than
fixed-horizon return labeling, because it incorporates the path
between t and t+H and encodes explicit risk management
[T4] (Source: Gemini, MiniMax).
Meta-labeling: a primary model emits a directional signal
(long/short/flat); a secondary "meta" model emits a binary decision of
whether to act on it, trained on the primary model's
predictions against realized outcomes. This permits dynamic position
sizing (act only above a confidence threshold) and lets the meta-model
consume side information the primary model does not see
[T4] (Source: Gemini, MiniMax).
Peer-reviewed support: Karakunnel et al. (2025),
"Algorithmic crypto trading using information-driven bars, triple
barrier labeling and deep learning," Financial Innovation,
doi:10.1186/s40854-025-00866-w [T1]; Karasan et al. (2024),
"Enhanced Genetic-Algorithm-Driven Triple Barrier Labeling Method and
Machine Learning Approach for Pair Trading Strategy in Cryptocurrency
Markets," Mathematics 12(5), 780, doi:10.3390/math12050780
[T2] (Source: MiniMax).
MiniMax's epistemic caveat is the most important sentence in
this subsection and is carried forward unsoftened: the
peer-reviewed evidence establishes that these techniques are stable
in production; their claimed superiority over fixed-horizon
labeling is asserted in the original book and has not been independently
replicated in a peer-reviewed controlled comparison
[T3] (Source: MiniMax). Gemini presents
triple-barrier and meta-labeling without this caveat (Source:
Gemini). The technique is defensible on ergonomic grounds — it
encodes risk management into the label — but it cannot be claimed to
improve performance.
Python: mlfinlab is not installable from PyPI
(commercial). Implement in pandas/numpy with
vectorized barrier-crossing logic [T6] (Source:
MiniMax).
No. Not at any conventional confidence level, and not by a small margin. All three reports reach this conclusion independently, by three different routes, with no dissent (Source: Qwen, Gemini, MiniMax).
C.1 The binomial route (MiniMax). Test H₀: true win
probability p₀ = 0.5 against H₁: p₁ > 0.5, using an exact binomial
test at α with 80% power. The required sample sizes, relabelled
here to match what they actually are — MiniMax's table is
headed "two-sided" but every figure in it reproduces exactly under a
one-sided exact binomial test [T3]:
| p₀ | p₁ (edge) | α | n required (one-sided exact) | n required (two-sided exact) |
|---|---|---|---|---|
| 0.50 | 0.55 (+5 pp) | 0.05 | 620 | 786 |
| 0.50 | 0.55 (+5 pp) | 0.01 | 1,007 | — |
| 0.50 | 0.60 (+10 pp) | 0.05 | 158 | 199 |
| 0.50 | 0.60 (+10 pp) | 0.01 | 252 | — |
| 0.50 | 0.70 (+20 pp) | 0.05 | 37 | 49 |
| 0.50 | 0.70 (+20 pp) | 0.01 | 64 | — |
| 0.40 | 0.50 (+10 pp) | 0.05 | 158 | — |
| 0.40 | 0.50 (+10 pp) | 0.01 | 253 | — |
(Source: MiniMax; column labels and the two-sided column
corrected. MiniMax's executive summary separately reports 783 / 194 / 47
for the three α = 0.05 rows — those are the two-sided normal
approximation, not an independent estimate, and the discrepancy
with its own table is an unacknowledged one-sided/two-sided switch
rather than a disagreement. MiniMax also prints a Cohen (1988/1992)
analytic formula, doi:10.1037/0033-2909.112.1.155, which is the
two-independent-proportions sample size — the wrong
formula for a one-sample test against a fixed p₀ — and it generates
neither set of numbers. The formula and its accompanying claim that "the
two methods agree to within rounding" are dropped
[T6].)
C.2 The continuous-return route. For a strategy
measured on returns rather than win rate,
n = (z_{1−α/2} + z_{0.80})² · σ²/μ². At a daily Sharpe of
0.1: n = (1.96 + 0.84)² / 0.1² = 784 trading days ≈
3.1 years (arithmetic verified)
[T3] (Source: MiniMax).
C.3 The t-statistic route (Gemini), and the cleanest
statement of the problem. To reach t ≥ 3.0 over N = 90 trading
days requires a daily Sharpe of 3/√90 = 0.3162, which
annualizes to 0.3162 × √252 = 5.02
(arithmetic independently verified) [T3]. An
annualized Sharpe of 5.02 is effectively non-existent in unleveraged
retail-accessible asset classes. Therefore a 90-day USD 100
deployment cannot even in principle demonstrate statistical proof of
edge [T3] (Source: Gemini). This is the
single most decision-relevant number in Section 8, it is arithmetically
correct, and it requires no assumptions about win rates or independence.
Gemini attaches no citation or tier tag to this derivation; it is
carried as [T3] because it is an unrefereed but
independently verifiable calculation, not a peer-reviewed finding.
C.4 How many independent trials are actually available in 90
days? MiniMax's per-strategy-class enumeration
[T3] (Source: MiniMax):
| Strategy class | Nominal trials | Effective independent n | Verdict |
|---|---|---|---|
| Daily-rebalanced US equity, liquid names | ~63 trading days | 30–50 | Below 37 even under the most generous independence assumption — insufficient |
| Event contracts (one/day, binary settlement) | ≤ 63 | 9–63 (most settle in 1–7 days) | Insufficient at the +10 pp threshold; marginally sufficient at +20 pp |
| Options (one trade/day, single-leg, 0–7 DTE) | ≤ 63 | 30–50 | Insufficient |
| Single 90-day equity holding | exactly 1 | 1 | Catastrophically insufficient — the most common retail approach and the most underpowered |
| Crypto, 24/7 | ≤ 90 | 50–80 | Insufficient at +10 pp; marginally sufficient at +20 pp only under optimistic independence |
Running multiple parallel strategies does not rescue this. The
inferential sample would be the number of independent strategies, but
the capital is one stake, so the strategies are not independent
in dollar space. The relevant n is the number of distinct bets,
not the number of strategies [T3] (Source:
MiniMax).
C.5 The autocorrelation haircut. Nominal trial
counts overstate the effective sample. Two adjustments, both from
MiniMax [T3]:
n_eff ≈ n · (1 − ρ)/(1 + ρ). For daily equity, ρ ≈ 0 →
n_eff ≈ n. For options P&L with volatility clustering, ρ ≈ 0.10 →
n_eff ≈ 0.82n. For binary event-contract P&L, ρ ≈ 0 → n_eff ≈
n.n_eff = n / q̄, where q̄ is the mean block length. For daily
equity returns q̄ ≈ 5–20, so for n = 63, n_eff ≈ 3.2 to
12.6 (arithmetic corrected — MiniMax prints "n/10 to n/3,"
i.e. 6.3–21, which does not follow from q̄ ∈ [5, 20]).MiniMax quotes three mutually inconsistent ranges for the effective sample across its own document (3–20, 5–30, and "5-30"), and separately asserts both n_eff ≈ n and n_eff ≈ 3–20 for the same daily-equity case. The defensible resolved range is n_eff ≈ 3 to 30 for a 63-day series, depending on the dependence assumption — a range whose entire span sits below the 37 trials needed to detect even a 20-percentage-point edge. The conclusion does not depend on which end of the range you accept, which is why the internal inconsistency does not rescue the experiment (Source: MiniMax, corrected).
C.6 The plain answer. The largest plausible win-rate
edge a retail strategy might carry against a 50% null is on the order of
10–20 percentage points, corresponding to a trade-level
Sharpe of roughly 0.2–0.5 [T6]. Detecting +10 pp at α =
0.05 with 80% power requires 158 independent trials
(one-sided) or 199 (two-sided). Detecting +20 pp
requires 37 or 49. The maximum number
of independent trials a 90-day experiment can produce is approximately
63, and after the autocorrelation haircut the effective
sample is approximately 3 to 30 [T3].
The experiment is, by construction, underpowered to validate
a strategy of any plausible effect size at conventional confidence
levels (Source: MiniMax). Qwen reaches the same
conclusion in plain language and should be quoted for it: "The
statistical power of such a short experiment is extremely low, meaning
it cannot reliably validate a new strategy's skill versus luck.
Any live result from this experiment would be considered
statistical noise" [T4] (Source: Qwen).
Gemini: the experiment's "only defensible value is infrastructure and
execution validation" [T4] (Source: Gemini).
C.7 The institutional analogue. Even ten-year
professional track records fail this test. Andrikogiannopoulou, A., Li,
Y., & Palia, D. (2019), "Reassessing False Discoveries in Mutual
Fund Performance: Skill, Luck, or Lack of Power?", Journal of
Finance 74(5), 2667–2702, doi:10.1111/jofi.12784 [T1],
critiquing Barras, L., Scaillet, O., & Wermers, R. (2010),
Journal of Finance 65(1), 179–216,
doi:10.1111/j.1540-6261.2009.01527.x [T1], shows that the
multiple-testing correction itself consumed the statistical power,
leaving the test unable to separate skill from luck for any fund with a
t-statistic below 3.0. MiniMax's retail extrapolation — that a 90-day
experiment carries roughly 1/10th the statistical power of a
ten-year fund performance test — is offered without derivation
and is carried as [T6] (Source: MiniMax; the two
underlying citations are [T1] and correct).
C.8 The one condition under which the experiment could
conclude something, and why it does not apply. MiniMax's own
adversarial self-critique concedes the escape hatch honestly: a 90-day
experiment can validate a strategy if the alternative
hypothesis is a very large edge — at p₁ = 0.9, 5–10 trials
suffice. But no documented strategy exhibits a 90% win rate
over 90 days [T3] (Source: MiniMax). Relaxing α to
0.10, going one-sided, or substituting a Bayesian 0.80-posterior
threshold all lower the required n — and all equally lower the bar for
the competing explanation, which is luck [T3].
C.9 What a "successful" outcome would actually mean.
If the account reaches USD 200, at least nine mutually compatible
explanations remain live, and the experiment cannot discriminate among
them: (a) a true edge, (b) luck, (c) a data artifact, (d) a
point-in-time failure, (e) a survivorship-bias failure, (f) an
unrecognized regime, (g) slippage understatement, (h) cost
understatement, (i) a fat-tail draw that landed on the upside
[T3] (Source: MiniMax). Success is not evidence of
skill; it is one observation drawn from a distribution the experiment is
too small to characterize.
Note on an excluded calculation. MiniMax attempts to quantify this with a Bayesian posterior, concluding "P(skill | success) ≈ 0.50 — the experiment is barely informative," and repeats the figure in its machine-readable appendix. The calculation is wrong by approximately seventeen orders of magnitude and, separately, uses a likelihood corresponding to 63 consecutive wins rather than the observed data, so even a correctly evaluated version would answer a different question. The 0.50 posterior is excluded from this report. The directional conclusion — that a short experiment is weakly informative — survives on the power arithmetic of §C.1–C.6 and does not need this calculation (Source: MiniMax, excluded).
C.10 The operational consequence. Treat the 90 days
as a pilot, not a trial. The informative outputs are
the ones that generalize: realized slippage, realized fill rates,
realized point-in-time failures, realized cost basis in basis points,
and realized calibration of any probabilistic forecast. Those are
measurable in a sample of 63 and transfer to a longer-horizon
experiment. The dollar P&L of the experiment is not
generalizable [T3] (Source: MiniMax).
Qwen reaches the same reframing in general terms — the purpose is "not
to prove a strategy works, but to learn about the practical challenges
of execution and the true distribution of outcomes in a real-world
setting" [T4] (Source: Qwen).
Gemini goes further than either and supplies a structured three-item Epistemic Salvage Plan, which is the most operationally specific version of this reframing across the three reports and is reproduced here in full (Source: Gemini):
p_i against binary
settlement outcomes y_i across every trade, then evaluate
the reliability and resolution partitions under Murphy's Score
Decomposition [T6]. (Gemini tags Murphy's
decomposition [T1] in its body but the citation appears
nowhere in its 47-entry bibliography; downgraded to [T6]
here and flagged, though the decomposition itself is standard
forecast-verification apparatus.) This is the formal machinery
behind "realized calibration," and it is the one output of a 63-trade
sample that is genuinely well-powered — calibration is estimable from
far fewer observations than edge is.P_mid and live fills P_fill, and use it to
build a realistic micro-capital friction model [T3]. This
directly measures trap A.7 rather than assuming it away, and it is the
one measurement that transfers unchanged to a larger or longer
experiment.[T3].All three reports converge on the pilot-not-trial reframing, which is the strongest consensus in this section. Gemini's version is the one to implement, because it names the measurements; MiniMax's adds realized point-in-time failures and realized cost basis to the list.
C.11 Operational rules, in priority order. From
MiniMax, with corrections applied [T3] (Source:
MiniMax):
[T6] — but the arithmetic is internally consistent and the
direction is right.)S₀ = √(V[max Ŝ]) is wrong. Replaced with
the published form, with S₀ defined over the variance of the N
trials' Sharpe ratios. Gemini's independent description (deflating PSR
for N, γ₃, γ₄) corroborates the corrected version.n/q̄ for q̄ ∈ [5,20] as n/10–n/3. Resolved to n_eff ≈
3–30 for a 63-day series, with the arithmetic corrected to 3.2–12.6
under the block-bootstrap estimate. The conclusion is invariant
across the whole range.arch.bootstrap has provided first-class SPA,
StepM, and MCS classes for years.arch version pin. Gemini 7.2.0 vs
MiniMax ≥ 8.0. Irreconcilable and unverifiable from the sources;
both dropped. Package identity retained, version claim
not.purgedcv >= 0.1.3. MiniMax only,
and self-contradictory within MiniMax (released on the report's own
research date, yet simultaneously a mature mlfinlab
successor, yet missing the functions it is recommended for).
Downgraded to [T6]; not recommended without
independent verification.[T4]
book provenance.[T6]; the citation is retained for the mechanism,
not the magnitude.[T6].10.1214/aos/1176348899 and 10.1093/rfs/4.4.867
(MiniMax). Absent from MiniMax's own bibliography and mismatched to the
methods named. Dropped. The correction menu is retained
on method names.[T1] in its Epistemic Salvage Plan but the citation appears
nowhere in its 47-entry bibliography. Downgraded to
[T6] and the gap flagged; the method is retained
because it is standard forecast-verification apparatus and is
load-bearing for the calibration measurement in §C.10.Three deep-research reports were commissioned against the same CASINO prompt. On this cluster they diverge more sharply than on any other section of the master report.
Qwen delivered nothing. Across roughly 6,700 words
the report names zero Python packages, produces no Table E, and contains
exactly two passing references to Python as a concept — "the need to be
operationalized via a Python-based system" and "an investor could build
a Python system that ingests real-time odds from a prediction market."
No pandas, no numpy, no backtester, no scoring library. This cluster is
a total gap in Qwen, not a thin one [T6]. (Source: Qwen
digest §"Cluster 7 — ESSENTIALLY ABSENT") Every conflict
adjudicated below is therefore a two-source contest between Gemini and
MiniMax, and every majority-rule test in this section reduces to
agreement-or-not between exactly two witnesses. Readers should weight
the resulting table accordingly: nothing here carries three-source
corroboration.
MiniMax delivered the deepest coverage and is the primary
source for Table E. It supplies 53 rows (≈51 distinct projects
— two same-codebase pairs are double-listed), all ten required
categories as separately numbered sections, an 18-layer reference
architecture, a 16-row avoid list, and per-package version, release
date, SPDX license, repository, star count, open-issue count,
maintenance status, and layer assignment [T6]. (Source:
MiniMax cluster-07 digest §§37–47, Table E)
Gemini delivered 26 packages with the same attribute
schema plus a verified_via: "PyPI API" flag on every JSON
record, a 6-package avoid list, and a cleaner 8-layer reference
architecture that maps directly onto the ingestion→execution pipeline
this section is asked to produce [T6]. (Source: Gemini
JSON appendix python_libraries,
unmaintained_libraries; Table E)
Both surviving sources claim live PyPI JSON API and GitHub REST API
retrieval on 2026-08-01. Neither source captured a single API
response. No JSON excerpts, no HTTP status codes, no per-record
retrieval timestamps, no ETags — only URLs that would return
the claimed fields if queried [T6]. Both digests flag this
independently and reach the same recommendation: treat every version
string, release date, star count, and open-issue count in this cluster
as an unsourced assertion requiring independent re-verification before
publication. (Source: MiniMax digest §6.1; Gemini digest §"Table E …
treat every number in Table E as unverified and re-derive from
PyPI/GitHub before any of it reaches the master report")
Three specific integrity problems compound this:
gh api CLI — 101 repositories total — but the source list
contains ~61 GitHub URLs [T6]. It also asserts "every
repository's stargazers_count was retrieved" while shipping
four n/a star cells, one of them annotated "n/a (the
leader)," which is not a data value [T6].https://pypi.org/project/<name>/json is not a valid
PyPI JSON path (the correct form is
https://pypi.org/pypi/<name>/json), and two package
names in the source list are typo'd into non-existent packages
(py-ro-ppl, pyalgo-trade) [T5].
URLs that do not resolve cannot have been parsed.vectorbt → github.com/polakfl/vectorbt,
riskfolio-lib →
github.com/dcgerard/Riskfolio-Lib, scores →
github.com/nswbusiness/scores,
kalshi-python-async →
github.com/kalshi/kalshi-python-async [T6].
(Source: Gemini digest citation-integrity flags) Where Gemini
and MiniMax disagree on a repository owner, this pattern is decisive
against Gemini.What survives this audit is the package selection and the
architecture-layer assignments, which both sources get
substantially right and which corroborate each other. What does not
survive is the metadata precision. The practical instruction to
the principal: install the stack below, then run
pip index versions <name> or query
https://pypi.org/pypi/<name>/json yourself to pin
actual versions. The library choices are load-bearing; the digits are
not.
Tier tags in this section. [T5] marks a
claim where a primary vendor or project document was named and the claim
is corroborated across both surviving sources. [T6] marks
single-source assertions and every star/issue count without exception.
[T1]–[T4] are near-absent from this cluster by
its nature — a package registry is not peer-reviewed literature — and
appear only where a named published method underwrites a library's
existence.
Five packages carry the ingestion layer, and the assignment is uncontested because only MiniMax surveyed it seriously.
yfinance 1.5.2 remains the default for US equity and ETF
end-of-day OHLCV, dividends, and splits [T6]. Its
constraint is legal, not technical: Yahoo's terms license personal,
non-commercial use only, and automated extraction plus redistribution
requires review [T5]. At USD 100 personal scale nobody will
challenge it; the principal should nonetheless understand that the
single most-used ingestion package in retail quant is the one operating
furthest outside its provider's terms. Gemini independently flags
yfinance as the only data source in its entire audit with
tos_restricts_automation: true,
point_in_time: false, and
survivorship_bias_free: false [T5].
(Source: Gemini Table F, MiniMax §37) That triple-negative
matters more for backtest integrity (Section 8) than for library
selection.
pandas-datareader 0.11.1 is canonical for FRED, World
Bank, and OECD macro series [T6]. Its Yahoo path broke in
2020 and has not returned — use yfinance for Yahoo,
pandas-datareader for official statistical sources
[T5]. MiniMax marks it "active" on a 2026-06-24 release
while elsewhere describing the project as having "the first release in
over a year," which is exactly the shape of claim worth re-verifying
before depending on it. The fallback is trivial: vendor the FRED REST
calls through httpx directly.
ccxt 4.5.70 unifies 105+ crypto venues behind one API
[T6]. Note the commercial boundary: CCXT Pro (WebSocket
streaming) is a paid product; the open-source package is REST-only
[T5].
ib-async 2.1.0 is the only actively maintained Python
framework speaking Interactive Brokers' native protocol, and both
sources agree on the version [T5]. They disagree on the
repository — MiniMax says ib-api-reloaded/ib_async, Gemini
says erdewit/ib_async [T6]. MiniMax supplies
the mechanism: the package was renamed and the repository moved after
the original maintainer's death in early 2024, superseding the abandoned
ib_insync [T6]. Gemini's erdewit
URL is the pre-transfer location. Use
ib-api-reloaded/ib_async.
polygon-api-client 1.16.3 covers higher-fidelity US
equities and options chains [T6]. MiniMax additionally
asserts a Polygon→Massive corporate rebrand on 2025-10-30 and a new
canonical package massive 2.8.0. I have excluded
massive from Table E. Its entire existence rests
on a README-sourced rebrand claim, and the asserted rebrand date is
identical to the asserted release date of
polygon-api-client 1.16.3 — a coincidence tidy enough that
MiniMax's own digest flags it as "plausible-sounding package whose
existence I cannot corroborate" [T6]. If the rebrand is
real, massive is the forward-looking name; verify before
writing an import statement against it.
The named gap. MiniMax searched PyPI on 2026-08-01
and found no first-class Python client for the Kalshi public
API, directing the principal to write a thin httpx
client against Kalshi's documented REST endpoints at
docs.kalshi.com [T5]. Gemini's JSON claims
kalshi-python-async 3.25.0 exists with 310 stars at
github.com/kalshi/kalshi-python-async [T6].
I resolve against Gemini and exclude the package. An
affirmative search that found nothing beats an assertion whose
repository URL falls in the same set Gemini's own digest flags as
fabricated-looking, and Kalshi has no history of publishing a
first-party async Python SDK under that name. Given that Kalshi is a
top-ranked vehicle in Section 10's strategy table, this is a material
gap: the principal must budget for hand-rolling the client.
statsmodels 0.14.6 anchors the layer — ARIMA, VAR, VECM,
panel regression, the standard test battery, and
statsmodels.tsa.statespace for SARIMAX,
UnobservedComponents, and VARMAX [T5] (both sources agree
on the version). The state-space submodule is the decision point that
removes a dependency: it covers Kalman filtering and structural time
series with tighter integration than any standalone filter package,
which is why filterpy (dead since 2018) does not appear in
Table E and pykalman appears only as an optional extra
[T6].
arch 8.0.0 is canonical for GARCH-family volatility,
unit-root testing, and bootstrap inference [T6]. Both
sources list it; they disagree on version (8.0.0 vs 7.2.0) and license
(NCSA vs MIT). I resolve to MiniMax on both counts — it cites the
project's pyproject.toml license declaration explicitly,
and NCSA is a permissive non-copyleft license functionally equivalent to
MIT for this use, which explains how Gemini could report MIT without
being wrong in spirit [T5].
linearmodels 7.0 handles panel data (fixed effects,
random effects, Fama-MacBeth), instrumental variables (2SLS, LIML, GMM),
and asset-pricing factor models [T5] — version agreed
across both sources. Same maintainer as arch, same NCSA
license, and the license conflict resolves identically.
pmdarima 2.1.1 is the Python auto.arima,
and it is slowing — one release in the trailing 365
days, with its lead maintainer having redirected primary effort to
statsforecast [T6]. It stays in Table E as a
reference implementation. For new code, use
statsforecast.AutoARIMA, whose README claims a 20× speedup
over pmdarima and 1.5× over R's forecast
[T3] — a vendor benchmark, weight accordingly.
pykalman 0.11.2 survives recency but earns no place in
the reference architecture; statsmodels.tsa.statespace
covers the same ground with better integration [T6].
Install it only for the narrow LiNGAM-style extensions MiniMax
cites.
Pick one primary PPL. Running two is a maintenance tax with no analytical payoff at this scale.
pymc is the recommendation for offline Bayesian modeling
on a single CPU: NUTS, ADVI, full MCMC, on the PyTensor symbolic
compiler (the Theano/Aesara successor) [T6]. The version is
genuinely contested — MiniMax says 6.2.0 (2026-07-23), Gemini says
5.17.0 (2026-07-02). I carry MiniMax's 6.2.0 because it supplies a
mechanism (6.x is a major API rewrite relative to 3.x, now stable)
rather than a bare number, but I flag it [T6]: a full
major-version jump asserted by one source against another's
contemporaneous 5.x claim is exactly the kind of divergence the
principal should resolve at pip install time. If 6.x is
real, expect breaking API changes against every PyMC tutorial written
before it.
arviz 1.2.0 is not a PPL — it is the diagnostics layer
for one. ESS, R̂, LOO, WAIC, trace plots, posterior predictive checks
[T6]. Install it alongside whichever sampler you choose,
always. A Bayesian forecast published without R̂ and ESS is an unaudited
number.
numpyro 0.21.0 (JAX backend) and cmdstanpy
1.3.0 (Stan interface) both carry two-source version agreement
[T5] — the strongest corroboration in this subsection. Take
numpyro if JAX is already resident on a GPU; take
cmdstanpy if you want Stan's HMC/NUTS implementation and
its documentation, which remains the best in the field. Neither is
necessary for a USD 100 experiment.
pyro-ppl 1.9.1 fails recency at ~26 months and is
excluded. Use numpyro, which shares its modeling idioms
[T6].
Four general-purpose frameworks all support probabilistic output
(predictive distributions or quantiles): sktime,
darts, statsforecast,
gluonts.
sktime 1.1.0 carries two-source version agreement
[T5] and is the highest-velocity framework in the survey —
a unified fit()/predict() API over 500+ models
spanning classical (ARIMA, ETS, Theta), ML (sklearn wrappers), deep
learning (N-BEATS, N-HiTS, TFT), and foundation models (Chronos,
TimesFM, PatchTST-FM) [T6].
darts 0.46.1 is comparable in scope with stronger neural
and probabilistic support, at the cost of a PyTorch-Lightning dependency
and a large install [T6]. Pick one of
sktime or darts as the primary API.
Running both against the same problem doubles the surface area for no
gain.
statsforecast 2.1.1 is the performance leader for the
statistical subset — Numba-compiled AutoARIMA, AutoETS, MSTL, Theta, CES
[T6]. It is the right default at this scale: lighter than
darts, faster than statsmodels on the same
models.
hierarchicalforecast 1.5.1 is the only library
in the surveyed universe that addresses hierarchical
reconciliation — BottomUp, TopDown, MinTrace, ERM, PERMBU, and
conformal methods [T6]. Single-source, but it answers a
required capability nothing else covers, so it stays in Table E with the
caveat attached.
neuralforecast 3.2.0 completes the Nixtla stack (NHITS,
NBEATS, TFT, PatchTST, TimesNet under one API) [T6]. Not
needed for the doubling experiment; included for forward extension.
gluonts 0.17.0 is AWS's probabilistic time-series
library and the reference implementation of DeepAR [T6].
Its cadence is slowing. Reach for darts first.
This is the highest-stakes subsection in the cluster, because the CASINO prompt's own framing is right: the reader will otherwise adopt a backtester on reputation. Reputation in this corner of the ecosystem is a lagging indicator by four to six years.
The four engines that pass.
vectorbt 1.1.0 — vectorized parameter sweeps over OHLCV
arrays, Numba-compiled [T5] (two-source version agreement).
Its license is the conflict: MiniMax reports Apache-2.0 with the
Commons Clause addendum, Gemini reports bare Apache-2.0
[T6]. I resolve to MiniMax — the Commons Clause claim is
specific, falsifiable, and matches the project's documented licensing
posture, whereas "Apache-2.0" is the answer you get by reading only the
SPDX field GitHub returns. Practical effect: source is publicly
available and free for individuals and organizations, but you may not
sell a product or service whose value derives primarily from the
software. Irrelevant for a personal experiment; material if the
principal ever packages the stack.
backtesting.py 0.6.6 — event-driven, single-instrument,
Bokeh plotting, fast research loop, single-author maintenance
[T6]. AGPL-3.0-or-later, which is genuine
copyleft: distributing a derivative — including over a network —
triggers source-disclosure obligations [T5]. Non-issue
privately; disqualifying for a hosted service.
nautilus-trader 1.230.0 — the most production-grade
option, Rust core, true event-driven order semantics, live and
backtest from the same code path, LGPL-3.0+ [T6]. Steepest
learning curve in the survey. Gemini and MiniMax disagree on version
(1.218.0 vs 1.230.0); given the project's multiple-releases-per-month
cadence, both are plausible snapshots and I carry the higher.
bt 1.2.0 — tree-structured portfolio composition
[T6]. Narrow but clean if you are composing weighted
sleeves rather than trading signals.
The avoid list — the part that actually protects the reader.
zipline-reloaded is the single most important resolution
in this section. MiniMax recommends it as "the live
replacement" for zipline, places it in its recommended
table, and marks it "active (slowing)" — while simultaneously stating
its own mechanically-enforced 365-day rule and reporting a release date
of 2025-07-19, which is 378 days before the 2026-08-01
research date [T6]. The package fails MiniMax's own test
and MiniMax did not notice. Gemini independently places
zipline/zipline-reloaded on its avoid list,
citing Cython compilation failures and a hard pin on Pandas <2.0
[T6]. Resolution: zipline-reloaded
moves to the unmaintained list. Both the arithmetic and the
second source agree; only MiniMax's unexamined prose dissents. Given
that the reference architecture assumes pandas 3.x, a hard
pandas<2.0 pin is independently disqualifying.
backtrader is the reputational trap the prompt
anticipated. Roughly 22,660 stars, ubiquitous in tutorials and YouTube
courses, last release 2023-04-19 (MiniMax) or
2023-04-08 (Gemini), repository dormant since 2024-08 [T4].
Treat it as a frozen codebase, not a maintained library. MiniMax also
understates its dormancy by roughly half ("12 months" for ~23.5 months),
one of seven arithmetic errors its own digest catches in the report's
dormancy statements [T6].
pyfolio and empyrical — both abandoned
Quantopian lineage [T6]. The commonly-recommended successor
pyfolio-reloaded was not found on PyPI as of
2026-08-01 (git repository only), so do not treat it as a
drop-in. Use skfolio's built-in performance summary for
tear-sheet output. MiniMax additionally recommends
quantstats here — but supplies no version, no release date,
no license, no repository, and no star count for it, in a report whose
sole purpose is per-package metadata. I have excluded
quantstats from Table E on exactly that basis
[T6].
pyalgotrade (2018), pybacktest (a single
2025 release against 4 stars and no community), and opstrat
(2021, Gemini-only) complete the list.
cvxpy 1.9.2 is the convex-programming DSL underneath
both PyPortfolioOpt and
Riskfolio-Lib [T5] (two-source version
agreement). Install it directly — the moment you need a custom objective
or constraint, you are writing cvxpy anyway.
Riskfolio-Lib 7.3.0 is the most capable single library
in the survey: 26 convex risk measures (variance, MAD, GMD, CVaR, EVaR,
CDaR, EDaR, RLVaR, Worst Realization), risk-parity variants,
hierarchical clustering (HRP, HERC), nested clustered optimization,
Worst-Case Mean-Variance, OWA, MVSK, Black-Litterman, entropy pooling,
cardinality constraints [T6]. Version agreed across both
sources [T5].
skfolio 0.20.1 is the integration point I
recommend despite being single-source. It is the only portfolio
library following the scikit-learn fit/predict
contract, and it ships CombinatorialPurgedCV and
WalkForward as first-class cross-validators in
skfolio.model_selection [T5] — MiniMax's
best-sourced API claim in the entire cluster, naming the package's
__init__ exports and showing the import statement. That
single fact collapses two requirements (portfolio optimization and
finance-appropriate validation) into one dependency and eliminates the
need for mlfinpy entirely.
PyPortfolioOpt 1.6.0 remains the most widely cited
mean-variance and Black-Litterman implementation, at a slowing cadence
[T6]. Keep it only if you want its specific Black-Litterman
API.
On Kelly. Neither source found a maintained
standalone Kelly library worth recommending, and that is the correct
outcome: the Kelly fraction f* = (bp − q)/b and its
fractional variants are a handful of lines of numpy, and
the higher-level optimizers expose Kelly objectives inside their
mean-risk layers [T6]. MiniMax routes the principal to
Riskfolio-Lib.optimization.mean_risk.portfolio_kelly —
do not trust that path. It is stated with
fully-qualified precision, zero documentation links, and no version pin,
and MiniMax's own digest flags it as "the single most actionable line in
§42 and the least sourced" [T6]. Implement Kelly directly
in numpy, where you can read every term.
The methodological caveat matters more than the API: Section 3
establishes that Kelly maximizes E[log W] over repeated
independent bets, which is not the objective here. A
fixed-multiple target under a hard deadline is a first-passage problem,
and under Dubins–Savage a subfair game rewards bold play, not
growth-optimal fractional sizing. Sizing code is cheap; using the right
objective function is the hard part.
vollib 1.0.11 is the recommendation for vanilla work:
Black, Black-Scholes, and Black-Scholes-Merton analytic prices; the full
standard Greek set (delta, gamma, vega, theta, rho, vanna, charm,
vomma); and implied volatility via Peter Jäckel's Let's Be
Rational algorithm [T2], which is the correct choice —
it is essentially machine-precision and non-iterative, and it matters
when you are inverting thousands of quotes to build a surface. MIT
licensed [T6].
Use vollib, not py_vollib.
MiniMax reports a 2026-06-01 rebrand in which py_vollib
became a deprecated transitional alias depending on vollib
for the implementation [T6]. Two anomalies deserve
flagging: the deprecated shim carries version 1.0.12 against the
canonical package's 1.0.11, which is backwards, and both are given the
identical release date. At least one of those version numbers is wrong.
The direction of the rename is what matters operationally, and
it is consistent across every mention. I have placed
py_vollib on the deprecated list rather than in Table
E.
pyfeng 0.5.0 covers SABR, Heston, NSVh, Schöbel-Zhu,
rough Heston, and multi-asset models [T6].
GPL-2.0 — the strictest copyleft in this cluster.
Academic-grade and entirely appropriate for private research; do not
distribute derived code without understanding the obligation.
QuantLib (Python bindings) is the serious term-structure
and exotics engine [T6]. Version is contested (MiniMax
1.43, Gemini 1.35) and I carry the higher. Repository: use
lballabio/QuantLib and
lballabio/QuantLib-SWIG. MiniMax presents
lballabio as a mirror of a
quantlib/QuantLib canonical repo, which inverts the actual
relationship; Gemini's lballabio/QuantLib-SWIG is correct
for the Python bindings specifically [T5]. This is the one
repository-attribution conflict Gemini wins outright.
FinancePy 1.0.1 covers swaps, bonds, FRAs, futures,
options, and FX at a slowing cadence, under
GPL-3.0-or-later [T6]. Overkill for a
pure-options track. Note MiniMax misspells the author as "Dominik
O'Kane" (correct: Dominic O'Kane).
At USD 100 scale, vollib alone is sufficient. Add
pyfeng only when you need stochastic volatility, and
QuantLib only when you need a real term structure.
This is the layer that converts a trading experiment into a measurable one, and it is where the meteorological literature imported in Section 6 becomes executable code.
scores is the recommendation — the most comprehensive
coverage of the meteorological metric set available in Python: Brier and
threshold-Brier, CRPS, FIRM, SEEPS, MAE/MSE/RMSE, Kling-Gupta
Efficiency, NSE, Flip-Flop Index, the Diebold-Mariano
test, Fractions Skill Score, and isotonic regression for
reliability diagrams [T6]. Both sources recommend it and
both assign it to the evaluation layer — the cleanest cross-source
agreement in the cluster on role, if not on version (MiniMax
2.6.0, Gemini 2.5.0) or repository (MiniMax nci/scores,
correct; Gemini nswbusiness/scores, flagged as
fabricated-looking by its own digest). Carry MiniMax on both.
scoringrules 0.11.0 (Zanetta & Allen) implements
CRPS, energy, variogram, interval, and quantile scores across NumPy,
JAX, PyTorch, and TensorFlow backends [T6]. Take it when
CRPS evaluation is in an inner loop and speed matters.
xskillscore 0.0.29 is the right tool only if forecasts
live in xarray [T6].
properscoring is dead — last release
2015, over a decade [T6]. Both sources flag it. It is still
the top result for "python proper scoring rules" in most searches, which
is precisely the failure mode this section exists to prevent.
One caution on API paths: MiniMax's worked recipe calls
scores.brier_score and scores.crps_ensemble as
top-level functions, but the package organizes metrics under submodules
and no documentation reference is given [T6]. Read the
module layout before writing imports.
scikit-learn 1.9.0 is table stakes, and both sources
agree on the version to within one day of release date — the strongest
metadata corroboration anywhere in this cluster [T5]. It
supplies TimeSeriesSplit for expanding-window walk-forward.
It does not implement purged or embargoed
cross-validation, and this is the gap that destroys most retail
ML backtests: overlapping labels leak across naive k-fold boundaries and
produce out-of-sample collapse [T2] (López de Prado, 2018).
(Source: Gemini ineffective_strategies entry
6)
skfolio.model_selection.CombinatorialPurgedCV
and .WalkForward are the answer [T5].
This is the single most useful finding in the cluster: a maintained,
BSD-3-licensed, sklearn-API-compatible implementation of combinatorial
purged cross-validation, in a package the principal already needs for
portfolio optimization. Import as
from skfolio.model_selection import CombinatorialPurgedCV, WalkForward.
Everything else in this space is a trap:
mlfinlab (Hudson & Thames) — no
PyPI release at all. The 4.9k-star public GitHub repository contains 11
commits, no tagged releases, and a README declaring "all rights
reserved"; the real code ships under a paid commercial license
[T6]. It is the most-cited López de Prado implementation
and it is not open-source software.mlfinpy 0.1.2 — the open-source
successor, last released 2024-10-09, 661 days before
the research date, the only release since inception, repository dormant
since 2025-01-23 [T6]. Fails recency decisively.timeseriescv 0.2 (2018) and
filterpy 1.4.5 (2018) — both dead, both
still widely linked [T6].Gemini contributes the gradient-boosting tier that MiniMax omits
entirely: lightgbm 4.7.0, xgboost 3.3.0,
catboost 1.2.8 [T6]. All three are
single-source with unverifiable metadata, but their existence and
maintenance status are not seriously in doubt, and a finance ML layer
without a boosted-tree implementation is incomplete. They enter Table E
flagged single-source. Wrap them in skfolio's CV splitters,
never in sklearn.model_selection.KFold.
pandera 0.32.1 validates DataFrame and Series schemas,
and its 0.32 line adds a Narwhals-powered backend spanning Polars,
pandas, Dask, and PySpark under one API [T6]. Gate
every ingested DataFrame behind a pandera schema before it
reaches the analytical pipeline. This is the operational form
of the data-integrity requirements Section 8 derives from the
backtest-overfitting literature — a schema check that fires on a
silently changed yfinance column layout is worth more than
any amount of downstream defensive coding.
pydantic validates records and contracts — API
responses, configuration, trade records, model artifacts
[T6]. Both sources include it; versions conflict (MiniMax
2.13.4, Gemini 2.12.5) with an internally inconsistent date ordering
between them. Carry the higher version, take Gemini's star count to fill
MiniMax's n/a cell, tag [T6].
prefect and dagster are the two credible
orchestrators. Pick one. Prefect wraps imperative
Python functions; Dagster models asset-centric DAGs [T6].
Both sources independently select prefect for the reference
architecture, so that is the recommendation — but the deciding factor at
this scale is which mental model fits, not which is better.
apache-airflow is too heavyweight for a personal stack, and
great-expectations has moved to a commercial "GX Core" with
its open-source edition in maintenance-only status; pandera
replaces it [T6].
mlflow 3.15.0 carries two-source version agreement
[T5] and is the canonical experiment tracker. The local
filesystem backend is entirely sufficient here. Log every backtest run —
parameters, metrics, artifacts — because the only defensible output of a
90-day USD 100 experiment is a complete, honest record of what was
tried.
optuna 4.9.0 handles hyperparameter search
[T6]. Use it sparingly and inside purged CV. An HPO loop
over a short financial time series is a machine for manufacturing
overfit Sharpe ratios; Section 8's multiple-testing corrections apply to
every trial it runs.
The two surviving sources propose different decompositions. MiniMax gives 18 numbered layers (0–17), which is granular but conflates substrate with pipeline stage. Gemini gives 8 layers that map cleanly onto the ingestion→execution flow. I adopt Gemini's 8-stage spine and populate it with MiniMax's package assignments, which are more complete and better differentiated. This resolves the structural conflict without losing content: MiniMax's substrate layers (0–2) become ambient dependencies rather than pipeline stages, which is what they are.
| # | Stage | Packages | Rationale |
|---|---|---|---|
| — | Substrate (ambient) | numpy, scipy, pandas,
polars, pyarrow |
Required by everything downstream. pandas for
modeling-API compatibility, polars for ingestion and
transformation speed, pyarrow for zero-copy interchange.
Not a pipeline stage. |
| 1 | Ingestion | yfinance (equity/ETF EOD),
pandas-datareader (FRED/World Bank), ccxt
(crypto), ib-async (IBKR/ForecastEx),
polygon-api-client (options chains), hand-rolled
httpx client (Kalshi) |
One adapter per venue, each returning a raw frame. The Kalshi client
must be written by hand — no maintained package exists
[T5]. |
| 2 | Validation | pandera (DataFrames), pydantic (API
records, config, trade objects) |
Every ingestion endpoint gated by a schema; every API response gated by a model. Failures surface here, not three layers down. |
| 3 | Storage | duckdb over partitioned Parquet, written via
pyarrow |
Analytical SQL on local files. No database server, no cloud
dependency, no cost. Both sources converge on this pairing independently
[T5]. |
| 4 | Feature computation | polars expressions + numpy +
statsmodels, arch (GARCH-derived features),
linearmodels (factor exposures) |
No feature store at this scale — compute lazily over the relevant
slice. Rolling statistics are one-line polars
expressions. |
| 5 | Forecasting | statsforecast (fast statistical baseline) →
sktime or darts (unified
framework) → pymc + arviz (Bayesian, where
posterior uncertainty is the point) → hierarchicalforecast
(if forecasts must reconcile across a grouping structure) |
Always fit the cheap baseline first. A neural forecaster that cannot beat AutoETS is a negative result worth recording. |
| 6 | Evaluation | scores (Brier, CRPS, reliability diagrams,
Diebold-Mariano), scoringrules (fast CRPS),
skfolio.model_selection.CombinatorialPurgedCV
(out-of-sample estimation), scikit-learn (calibration
curves) |
Proper scoring rules on every probabilistic forecast; CPCV on every strategy evaluation. Deflated Sharpe computed from the per-path Sharpe distribution CPCV produces. |
| 7 | Sizing | numpy (Kelly and fractional-Kelly, implemented
directly), cvxpy (custom constrained objectives),
skfolio (sklearn-API portfolio construction),
Riskfolio-Lib (advanced risk measures when needed) |
Kelly is ten lines you should be able to read. Reach for the optimizers when constraints get real. |
| 8 | Execution decision | backtesting.py (single-strategy research loop) and
vectorbt (parameter sweeps) for simulation;
nautilus-trader for event-driven semantics and any live
path; vollib / pyfeng for options pricing and
Greeks at decision time |
Gemini assigns nautilus_trader to execution, MiniMax to
backtesting. Both are right: it is the only engine here that runs the
same code in backtest and live, which is exactly why it belongs at the
decision boundary. |
| 9 | Logging & orchestration | mlflow (every run: parameters, metrics, artifacts),
prefect (nightly flow: ingest → validate → recompute →
evaluate → log), optuna (tuning, inside purged CV
only) |
The experiment's only durable output. Log the losers as carefully as the winners — Section 10's epistemic salvage plan depends entirely on this layer being honest. |
On the footprint claim. MiniMax asserts three
separate times that this architecture runs on "19 actively maintained
packages plus four core computational dependencies." Counting distinct
package names across its own 18-layer table yields roughly
35 packages, or ~30 after removing the four named core
dependencies [T6]. The "19 packages" figure understates the
real dependency surface by about 50%, and it is precisely the kind of
tractable-sounding number a reader lifts verbatim. The architecture
above is honest about its size: roughly 30 packages if you install every
stage, and closer to 18 if you take one option per fork (one forecasting
framework, one backtester, one orchestrator, skip the Bayesian and
hierarchical layers).
The closing judgment from MiniMax is the right one to carry forward:
"a perfect software stack is necessary but not sufficient for the
doubling experiment; the literature surveyed in clusters 3, 5, and 6 is
the binding constraint on the probability of success."
[T6] No amount of tooling converts a negative-expectation
problem into a positive one. What the stack buys is the ability to
measure the outcome honestly — which, per Section 10, is the
only defensible deliverable this experiment can produce.
All star and open-issue counts are [T6] without
exception — neither source evidenced a single API response. Version and
release-date columns are [T5] where both surviving sources
agree on the version and [T6] where a single source asserts
it or the two conflict. Provenance: MM = MiniMax,
G = Gemini, MM+G = both. Qwen
contributed zero rows.
| # | Package | Version | Latest release | License | Stars | Open issues | Maintenance status | Architecture layer | Source URL consulted | Provenance |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | numpy |
2.5.1 | 2026-07-04 | BSD-3-Clause | 32,469 | 2,317 | active | Substrate | pypi.org/pypi/numpy/json ·
github.com/numpy/numpy |
MM |
| 2 | scipy |
1.18.0 | 2026-06-19 | BSD-3-Clause | 14,875 | 1,846 | active | Substrate | pypi.org/pypi/scipy/json ·
github.com/scipy/scipy |
MM |
| 3 | pandas |
3.0.5 | 2026-07-22 | BSD-3-Clause | 49,388 | 2,924 | active | Substrate | pypi.org/pypi/pandas/json ·
github.com/pandas-dev/pandas |
MM |
| 4 | polars |
1.43.2 | 2026-08-01 [T6] |
MIT | 39,156 | 2,846 | active | Substrate / feature computation | pypi.org/pypi/polars/json ·
github.com/pola-rs/polars |
MM+G (version conflict) |
| 5 | pyarrow |
25.0.0 | 2026-07-10 | Apache-2.0 | 16,969 | 2,554 | active | Substrate / columnar interchange | pypi.org/pypi/pyarrow/json ·
github.com/apache/arrow |
MM |
| 6 | duckdb |
1.5.5 | 2026-07-22 | MIT | 24,500 (main repo) | 420 | active | Storage | pypi.org/pypi/duckdb/json ·
github.com/duckdb/duckdb |
MM+G (version conflict) |
| 7 | yfinance |
1.5.2 | 2026-07-23 | Apache-2.0 | 24,856 | 169 | active (high) | Ingestion — Yahoo EOD/OHLCV | pypi.org/pypi/yfinance/json ·
github.com/ranaroussi/yfinance |
MM |
| 8 | pandas-datareader |
0.11.1 | 2026-06-24 [T6] |
BSD-3-Clause | 3,226 | 145 | active (slowing) | Ingestion — FRED / World Bank / OECD | pypi.org/pypi/pandas-datareader/json ·
github.com/pydata/pandas-datareader |
MM |
| 9 | ccxt |
4.5.70 | 2026-07-29 | MIT | 43,470 | 937 | active (very high) | Ingestion — crypto CEX/DEX | pypi.org/pypi/ccxt/json ·
github.com/ccxt/ccxt |
MM |
| 10 | ib-async |
2.1.0 [T5] |
2025-12-08 [T6] |
BSD-2-Clause [T6] |
1,707 | 89 | active | Ingestion + live execution — IBKR / ForecastEx | pypi.org/pypi/ib-async/json ·
github.com/ib-api-reloaded/ib_async |
MM+G (repo + license conflict) |
| 11 | polygon-api-client |
1.16.3 | 2025-10-30 | MIT | 1,490 | 19 | active | Ingestion — US equities / options chains | pypi.org/pypi/polygon-api-client/json ·
github.com/polygon-io/client-python |
MM |
| 12 | statsmodels |
0.14.6 [T5] |
2025-12-05 [T6] |
BSD-3-Clause | 11,546 | 2,888 | active | Econometrics + state space | pypi.org/pypi/statsmodels/json ·
github.com/statsmodels/statsmodels |
MM+G (date conflict) |
| 13 | arch |
8.0.0 [T6] |
2025-10-21 [T6] |
NCSA [T5] |
1,548 | 51 | active | GARCH / volatility / unit root | pypi.org/pypi/arch/json ·
github.com/bashtage/arch |
MM+G (version + license conflict) |
| 14 | linearmodels |
7.0 [T5] |
2025-10-21 [T6] |
NCSA [T5] |
1,060 | 56 | active | Panel / IV / asset pricing | pypi.org/pypi/linearmodels/json ·
github.com/bashtage/linearmodels |
MM+G (license conflict) |
| 15 | pmdarima |
2.1.1 | 2025-11-17 | MIT | 1,732 | 64 | slowing | Auto-ARIMA (reference only) | pypi.org/pypi/pmdarima/json ·
github.com/alkaline-ml/pmdarima |
MM |
| 16 | pykalman |
0.11.2 | 2026-01-31 | BSD-3-Clause | 1,327 | 85 | active (slowing) | State space (optional) | pypi.org/pypi/pykalman/json ·
github.com/pykalman/pykalman |
MM |
| 17 | sktime |
1.1.0 [T5] |
2026-07-28 [T6] |
BSD-3-Clause | 9,896 | 2,371 | active | Forecasting framework | pypi.org/pypi/sktime/json ·
github.com/sktime/sktime |
MM+G |
| 18 | statsforecast |
2.1.1 [T6] |
2026-07-16 | Apache-2.0 | 4,854 | 139 | active | Forecasting — fast statistical | pypi.org/pypi/statsforecast/json ·
github.com/Nixtla/statsforecast |
MM+G (version conflict) |
| 19 | darts |
0.46.1 [T6] |
2026-07-20 | Apache-2.0 | 9,480 | 215 | active | Forecasting — unified + neural | pypi.org/pypi/darts/json ·
github.com/unit8co/darts |
MM+G (version conflict) |
| 20 | gluonts |
0.17.0 | 2026-07-31 | Apache-2.0 | 5,221 | 470 | active (slowing) | Forecasting — probabilistic (DeepAR) | pypi.org/pypi/gluonts/json ·
github.com/awslabs/gluonts |
MM |
| 21 | neuralforecast |
3.2.0 | 2026-07-10 | Apache-2.0 | n/a | n/a | active | Forecasting — deep learning | pypi.org/pypi/neuralforecast/json ·
github.com/Nixtla/neuralforecast |
MM |
| 22 | hierarchicalforecast |
1.5.1 | 2026-03-04 | Apache-2.0 | 752 | 7 | active | Forecasting — hierarchical reconciliation | pypi.org/pypi/hierarchicalforecast/json ·
github.com/Nixtla/hierarchicalforecast |
MM |
| 23 | pymc |
6.2.0 [T6] |
2026-07-23 [T6] |
Apache-2.0 | 9,695 | 479 | active | Bayesian inference (primary PPL) | pypi.org/pypi/pymc/json ·
github.com/pymc-devs/pymc |
MM+G (major-version conflict) |
| 24 | pytensor |
3.2.3 | 2026-07-25 | BSD-3-Clause | n/a | n/a | active | Bayesian — symbolic compiler | pypi.org/pypi/pytensor/json ·
github.com/pymc-devs/pytensor |
MM |
| 25 | arviz |
1.2.0 | 2026-06-12 | Apache-2.0 | n/a | n/a | active | Bayesian — posterior diagnostics | pypi.org/pypi/arviz/json ·
github.com/arviz-devs/arviz |
MM |
| 26 | numpyro |
0.21.0 [T5] |
2026-05-02 [T6] |
Apache-2.0 | 2,730 | 68 | active | Bayesian — JAX / GPU PPL | pypi.org/pypi/numpyro/json ·
github.com/pyro-ppl/numpyro |
MM+G |
| 27 | cmdstanpy |
1.3.0 [T5] |
2025-10-20 [T6] |
BSD-3-Clause | 198 | 29 | active | Bayesian — Stan HMC/NUTS | pypi.org/pypi/cmdstanpy/json ·
github.com/stan-dev/cmdstanpy |
MM+G |
| 28 | vectorbt |
1.1.0 [T5] |
2026-07-05 [T6] |
Apache-2.0 + Commons Clause [T5] |
8,515 | 136 | active | Backtesting — vectorized sweeps | pypi.org/pypi/vectorbt/json ·
github.com/polakowo/vectorbt |
MM+G (license conflict) |
| 29 | backtesting (backtesting.py) |
0.6.6 | 2026-07-22 | AGPL-3.0-or-later | 8,745 | 61 | active (single maintainer) | Backtesting — event-driven research | pypi.org/pypi/backtesting/json ·
github.com/kernc/backtesting.py |
MM |
| 30 | bt |
1.2.0 | 2026-04-25 | MIT | 2,954 | 83 | active | Backtesting — tree/portfolio composition | pypi.org/pypi/bt/json ·
github.com/pmorissette/bt |
MM |
| 31 | nautilus-trader |
1.230.0 [T6] |
2026-06-29 [T6] |
LGPL-3.0-or-later [T5] |
25,180 | 82 | active (very high) | Execution decision — event-driven, Rust core | pypi.org/pypi/nautilus-trader/json ·
github.com/nautechsystems/nautilus_trader |
MM+G (version conflict) |
| 32 | cvxpy |
1.9.2 [T5] |
2026-06-22 [T6] |
Apache-2.0 | 6,293 | 192 | active | Sizing — convex programming | pypi.org/pypi/cvxpy/json ·
github.com/cvxpy/cvxpy |
MM+G |
| 33 | PyPortfolioOpt |
1.6.0 | 2026-02-26 | MIT | 5,922 | 109 | active (slowing) | Sizing — mean-variance / Black-Litterman | pypi.org/pypi/pyportfolioopt/json ·
github.com/robertmartin8/PyPortfolioOpt |
MM |
| 34 | Riskfolio-Lib |
7.3.0 [T5] |
2026-05-31 [T6] |
BSD-3-Clause | 4,420 | 28 | active | Sizing — advanced risk measures | pypi.org/pypi/riskfolio-lib/json ·
github.com/dcajasn/Riskfolio-Lib |
MM+G |
| 35 | skfolio |
0.20.1 | 2026-04-21 | BSD-3-Clause | 2,085 | 21 | active | Sizing + CombinatorialPurgedCV / WalkForward
[T5] |
pypi.org/pypi/skfolio/json ·
github.com/skfolio/skfolio |
MM |
| 36 | vollib |
1.0.11 | 2026-06-01 | MIT | 420 | 1 | active (single maintainer) | Options — BSM price / Greeks / IV | pypi.org/pypi/vollib/json ·
github.com/vollib/py_vollib |
MM |
| 37 | pyfeng |
0.5.0 | 2026-05-26 | GPL-2.0 | 184 | 2 | active (academic) | Options — SABR / Heston / rough vol | pypi.org/pypi/pyfeng/json ·
github.com/PyFE/PyFENG |
MM |
| 38 | QuantLib (Python bindings) |
1.43 [T6] |
2026-07-14 [T6] |
BSD-3-Clause (QuantLib modified) | 1,900 | 35 | active | Options — term structure / exotics | pypi.org/pypi/QuantLib/json ·
github.com/lballabio/QuantLib-SWIG |
MM+G (version + repo conflict) |
| 39 | FinancePy |
1.0.1 | 2025-08-31 | GPL-3.0-or-later | 3,080 | 55 | slowing | Options / rates / credit (optional) | pypi.org/pypi/financepy/json ·
github.com/domokane/FinancePy |
MM |
| 40 | scores |
2.6.0 [T6] |
2026-07-17 [T6] |
Apache-2.0 | 228 | 104 | active | Evaluation — Brier / CRPS / PIT / Diebold-Mariano | pypi.org/pypi/scores/json ·
github.com/nci/scores |
MM+G (version + repo conflict) |
| 41 | scoringrules |
0.11.0 | 2026-06-06 | Apache-2.0 | 97 | 17 | active | Evaluation — fast multi-backend CRPS | pypi.org/pypi/scoringrules/json ·
github.com/frazane/scoringrules |
MM |
| 42 | xskillscore |
0.0.29 | 2026-02-18 | Apache-2.0 | 242 | 52 | active (slowing) | Evaluation — xarray skill scores | pypi.org/pypi/xskillscore/json ·
github.com/xarray-contrib/xskillscore |
MM |
| 43 | scikit-learn |
1.9.0 [T5] |
2026-06-02 [T5] |
BSD-3-Clause | 66,849 | 2,109 | active | ML estimators + calibration + TimeSeriesSplit |
pypi.org/pypi/scikit-learn/json ·
github.com/scikit-learn/scikit-learn |
MM+G |
| 44 | lightgbm |
4.7.0 | 2026-05-04 | MIT | 16,500 | 340 | active | ML — gradient boosting | pypi.org/pypi/lightgbm/json ·
github.com/microsoft/LightGBM |
G |
| 45 | xgboost |
3.3.0 | 2026-06-20 | Apache-2.0 | 26,100 | 480 | active | ML — gradient boosting | pypi.org/pypi/xgboost/json ·
github.com/dmlc/xgboost |
G |
| 46 | catboost |
1.2.8 | 2026-04-29 | Apache-2.0 | 8,100 | 390 | active | ML — categorical gradient boosting | pypi.org/pypi/catboost/json ·
github.com/catboost/catboost |
G |
| 47 | pydantic |
2.13.4 [T6] |
2026-05-06 [T6] |
MIT | 22,400 | 210 | active | Validation — records / config / contracts | pypi.org/pypi/pydantic/json ·
github.com/pydantic/pydantic |
MM+G (version conflict) |
| 48 | pandera |
0.32.1 | 2026-06-29 | MIT | 4,413 | 448 | active | Validation — DataFrame schemas | pypi.org/pypi/pandera/json ·
github.com/unionai-oss/pandera |
MM |
| 49 | prefect |
3.8.1 [T6] |
2026-07-30 [T6] |
Apache-2.0 | 23,518 | 823 | active | Orchestration | pypi.org/pypi/prefect/json ·
github.com/PrefectHQ/prefect |
MM+G (version conflict) |
| 50 | dagster |
1.13.16 | 2026-07-30 | Apache-2.0 | 14,110 | 1,801 | active | Orchestration (alternative to Prefect) | pypi.org/pypi/dagster/json ·
github.com/dagster-io/dagster |
MM |
| 51 | mlflow |
3.15.0 [T5] |
2026-07-31 [T6] |
Apache-2.0 | 27,318 | 2,091 | active (very high) | Logging — experiment tracking | pypi.org/pypi/mlflow/json ·
github.com/mlflow/mlflow |
MM+G |
| 52 | optuna |
4.9.0 | 2026-06-01 | MIT | 14,591 | 18 | active | Logging — hyperparameter search | pypi.org/pypi/optuna/json ·
github.com/optuna/optuna |
MM |
52 rows, 52 distinct packages. Duplicate and alias
rows present in the source reports were collapsed: MiniMax's
massive/polygon-api-client pair and
py_vollib/vollib pair each shared one
repository and one set of star/issue counts, and MiniMax double-printed
neuralforecast in its §40 table alongside a placeholder
junk row (foundation-model-eval | (not surveyed) | n/a …)
that was drafting scaffolding rather than data.
Every package below appears in at least one source's exclusion list, or fails the 365-day recency test on the sources' own reported dates. Ordered by likelihood the principal encounters it in a tutorial.
| Package | Last release | Reason excluded | Replacement | Flagged by |
|---|---|---|---|---|
backtrader 1.9.78.123 |
2023-04-19 (MM) / 2023-04-08 (G) | ~23.5 months dormant; repo untouched since 2024-08. **~22,660 stars
make it the most-recommended dead backtester in Python**
[T4]. No native asyncio, no modern broker WebSocket
drivers. |
backtesting.py, vectorbt,
nautilus-trader |
MM + G |
zipline 1.4.1 |
2020-10-05 | Quantopian shut down late 2020; upstream has had no commits since. | nautilus-trader |
MM + G |
zipline-reloaded 3.1.1 |
2025-07-19 | 378 days before the research date — fails the 365-day
window. MiniMax recommended it anyway, in violation of its own
stated rule. Gemini independently flags Cython build failures and a hard
pandas<2.0 pin, incompatible with the
pandas 3.x substrate. |
nautilus-trader, backtesting.py |
MM (arithmetic) + G (explicit) |
pyalgotrade 0.20 |
2018-08-21 (MM) / 2018-08-02 (G) | ~8 years dormant; non-functional on Python 3.10+. Persists in legacy tutorials. | vectorbt, backtesting.py |
MM + G |
pybacktest 1.1.8 |
2025-03-27 | Technically inside the window, but the single 2025 release was the first since 2015 against 4 stars, 5 issues, no community, no documentation. Recency is necessary, not sufficient. | backtesting.py |
MM |
pyfolio 0.9.2 |
2019-04-15 (MM) / 2019-06-21 (G) | Quantopian lineage, abandoned ~7.3 years. Deprecated pandas/empyrical calls now raise at runtime. | skfolio Portfolio.summary() |
MM + G |
pyfolio-reloaded |
n/a | Community fork not present on PyPI as of 2026-08-01 — git repository only. Do not treat as a drop-in successor. | skfolio |
MM |
empyrical 0.5.5 |
2020-10-13 | Same Quantopian lineage; ~5.8 years since release. | skfolio, Riskfolio-Lib |
MM |
mlfinlab (Hudson & Thames) |
n/a — never on PyPI | Public 4.9k-star repo contains 11 commits, no tags, README declaring "all rights reserved." Real code is paid-commercial. Not open-source software. | skfolio.model_selection.CombinatorialPurgedCV |
MM |
mlfinpy 0.1.2 |
2024-10-09 | 661 days — the only release since project inception; repo dormant since 2025-01-23; 1 open issue against 79 stars. | skfolio.model_selection.CombinatorialPurgedCV |
MM |
timeseriescv 0.2 |
2018-09-07 | ~7.9 years dormant. Purged walk-forward CV now in
skfolio. |
skfolio.model_selection.WalkForward |
MM |
filterpy 1.4.5 |
2018-10-10 | ~7.8 years dormant. Foundational Kalman/EKF reference, no longer maintained. | statsmodels.tsa.statespace, pykalman |
MM |
properscoring 0.1 |
2015-11-12 (MM) / 2015-05-20 (G) | Over a decade. Still the top search result for Python proper scoring rules. | scores, scoringrules,
xskillscore |
MM + G |
uncertainty-toolbox |
n/a — not located on PyPI | Google research project, unmaintained since 2021. | scores reliability-diagram functions |
MM |
pyro-ppl 1.9.1 |
2024-06-02 | ~26 months; fails recency. | numpyro (same modeling idioms) |
MM |
py_vollib 1.0.12 |
2026-06-01 | Deprecated by its own maintainer — transitional
alias depending on vollib for the implementation. Passes
recency but is explicitly end-of-life. Note the anomaly: version 1.0.12
exceeds the canonical vollib 1.0.11 it wraps. |
vollib |
MM |
py-vollib-vectorized 0.1.1 |
2021-02-28 | ~5.4 years dormant; vollib now vectorizes
natively. |
vollib |
MM |
opstrat |
2021-07-14 | Unmaintained options-plotting script. | QuantLib, scipy, vollib |
G |
ffn |
2022-08-15 | ~4 years dormant. | Riskfolio-Lib, skfolio |
MM |
pyflux 0.9.1 |
2017 | ~9 years dormant. | statsmodels, pymc |
MM |
ib_insync |
abandoned (2024) | Original maintainer died early 2024; project renamed and transferred. | ib-async |
MM |
great-expectations (OSS) |
n/a | Project moved to commercial "GX Core"; open-source edition is maintenance-only. | pandera |
MM |
apache-airflow |
active | Not unmaintained — excluded for weight, not health. Disproportionate for a single-user stack. | prefect, dagster |
MM |
23 entries; 22 genuinely unmaintained, deprecated, or
unavailable (apache-airflow is listed for
completeness and is excluded on footprint grounds only).
Qwen omitted the entire cluster — every architecture layer, all ten required categories, Table E in full. No other gap in this section compares.
Relative to MiniMax, Gemini omitted six layers the other
source covered: data validation (pandera — Gemini
has pydantic for records but nothing for DataFrame
schemas), hierarchical forecasting (hierarchicalforecast),
finance-appropriate cross-validation as an explicit capability (Gemini
names López de Prado's purged-CV problem in its
ineffective-strategies list but assigns no library that solves
it), options Greeks and IV surfaces beyond QuantLib (no
vollib, no stochastic-vol library), hyperparameter
optimization (optuna), and macro/FRED ingestion
(pandas-datareader).
Relative to Gemini, MiniMax omitted the gradient-boosting
tier entirely — no lightgbm, xgboost,
or catboost anywhere in 53 rows, despite §45 nominally
covering "standard estimators (linear models, trees, gradient
boosting)." Those three rows enter Table E on Gemini's authority
alone.
Both sources omitted a news and text-ingestion layer (no NLP, no filings parser beyond raw EDGAR JSON), consistent with Section 12's finding that neither report located a free news data source.
Twenty-four conflicts adjudicated. Qwen contributed no Python content, so every contest is Gemini vs. MiniMax; no conflict was resolvable by three-source majority.
zipline-reloaded — active vs.
abandoned. MiniMax lists it as "active (slowing)" in its
recommended table and promotes it as "the live replacement" for
zipline; Gemini places it on its avoid list (Cython
failures, pandas<2.0 pin). Resolved:
unmaintained. MiniMax's own reported release date (2025-07-19)
is 378 days before its own 2026-08-01 research date, failing its own
mechanically-stated 365-day rule — its arithmetic contradicts its prose,
and Gemini independently agrees. The single most consequential
correction in this section.kalshi-python-async 3.25.0 — exists vs. does
not exist. Gemini asserts it with 310 stars at
github.com/kalshi/kalshi-python-async; MiniMax searched
PyPI on 2026-08-01 and reports no Kalshi Python client.
Resolved: excluded. An affirmative negative search
beats an assertion whose repo URL falls in the set Gemini's own digest
flags as fabricated-looking. Consequence recorded as a named gap: the
principal must hand-roll an httpx client.arch version and license. MiniMax
8.0.0 / NCSA vs. Gemini 7.2.0 / MIT. Resolved: 8.0.0,
NCSA. MiniMax cites the project's pyproject.toml
license declaration explicitly; NCSA is permissive and functionally
MIT-equivalent, which explains Gemini's approximation.linearmodels license. MiniMax NCSA vs.
Gemini MIT (version 7.0 agreed). Resolved: NCSA — same
maintainer and same declaration mechanism as arch;
consistency favors MiniMax.pymc major version. MiniMax 6.2.0 vs.
Gemini 5.17.0. Resolved: 6.2.0, tagged
[T6]. MiniMax supplies a mechanism (6.x is a
stabilized major API rewrite) rather than a bare number, but a full
major-version divergence between contemporaneous sources is unresolvable
from the digests and must be checked at install time.vectorbt license. MiniMax Apache-2.0
+ Commons Clause vs. Gemini bare Apache-2.0 (version
1.1.0 agreed). Resolved: Commons Clause included. The
specific, falsifiable claim beats the SPDX field GitHub returns, which
does not represent addenda.QuantLib version and repository.
MiniMax 1.43 at quantlib/QuantLib ("mirror at
lballabio/QuantLib") vs. Gemini 1.35 at
lballabio/QuantLib-SWIG. Resolved: version 1.43
(MiniMax), repository lballabio/QuantLib-SWIG
(Gemini). MiniMax inverts the canonical/mirror relationship —
lballabio is upstream — and Gemini correctly names the SWIG
bindings repo for the Python package. The only repository conflict
Gemini wins.scores version and repository. MiniMax
2.6.0 at nci/scores vs. Gemini 2.5.0 at
nswbusiness/scores. Resolved: MiniMax on
both. nci/scores is the plausible owner
(Australia's NCI); Gemini's owner is flagged fabricated-looking by its
own digest.ib-async repository and license.
MiniMax ib-api-reloaded/ib_async / BSD-2-Clause vs. Gemini
erdewit/ib_async / BSD-3-Clause (version 2.1.0 agreed).
Resolved: ib-api-reloaded/ib_async.
MiniMax supplies the mechanism — rename and transfer after the original
maintainer's death in early 2024. License left [T6];
neither source sourced it.duckdb star count. MiniMax 174
(annotated: duckdb/duckdb-python client repo) vs. Gemini
24,500 (duckdb/duckdb). Resolved: 24,500 against
the main project repo. MiniMax's number is correct for the
sub-repo it names but misleads a reader scanning a maintenance-health
column. Version 1.5.5 carried from MiniMax.statsforecast version. MiniMax 2.1.1
vs. Gemini 2.0.3. Resolved: 2.1.1 — higher and
plausible for a high-cadence Nixtla project; MiniMax's later stated date
is internally consistent with it.darts version. MiniMax 0.46.1 vs.
Gemini 0.31.0. Resolved: 0.46.1 — a 15-minor-version
gap is implausible as noise; the higher figure is the more likely
current one and MiniMax's date supports it.nautilus-trader version. MiniMax
1.230.0 (2026-06-29) vs. Gemini 1.218.0 (2026-07-28). Resolved:
1.230.0. Gemini pairs a lower version with a
later date, which is internally inconsistent; the project's
multiple-releases-per-month cadence makes 1.230 plausible by
mid-2026.prefect version. MiniMax 3.8.1 vs.
Gemini 3.4.8. Resolved: 3.8.1 — higher and consistent
with the project's release cadence.pydantic version. MiniMax 2.13.4
(2026-05-06) vs. Gemini 2.12.5 (2026-07-08). Resolved: 2.13.4,
tagged [T6]. Gemini again pairs a lower version
with a later date. Gemini's star count (22,400) was merged in to fill
MiniMax's n/a cell.polars version. MiniMax 1.43.2
(2026-08-01) vs. Gemini 1.21.0 (2026-07-22). Resolved: 1.43.2,
tagged [T6] with a caveat. MiniMax's date is the
same calendar day as its research date — a zero-day-old release is
exactly the shape of a value generated to satisfy a recency rule rather
than observed.statsmodels release date. MiniMax
2025-12-05 vs. Gemini 2026-04-10 (version 0.14.6 agreed by both).
Resolved: 2025-12-05, tagged [T6]. Same
version cannot have two release dates; MiniMax's date is corroborated by
its own narrative (prior release 0.14.5 in 2024-10, slow cadence). Both
pass recency either way, so nothing downstream turns on it.cmdstanpy / numpyro /
sktime / cvxpy / Riskfolio-Lib /
mlflow release dates. Versions agree across both
sources in every case; dates differ by weeks to months.
Resolved: MiniMax dates carried, versions upgraded to
[T5] on the strength of two-source version
agreement. No recency verdict changes under either source's
date.massive 2.8.0 — package existence.
MiniMax asserts it as the new canonical name after a claimed 2025-10-30
Polygon→Massive rebrand; no other source mentions it. Resolved:
excluded from Table E, noted in prose. The asserted rebrand
date is identical to the asserted release date of
polygon-api-client 1.16.3 — a coincidence MiniMax's own
digest flags as suggesting one date was reused.
polygon-api-client retained instead.quantstats — recommended with zero
metadata. MiniMax recommends it twice as the
pyfolio tear-sheet substitute but supplies no version,
date, license, repository, or star count anywhere in a report built
entirely on per-package metadata. Resolved: excluded.
skfolio's built-in performance summary substituted.py_vollib vs. vollib.
MiniMax lists both in Table E with py_vollib 1.0.12
above vollib 1.0.11 despite describing
py_vollib as the deprecated alias that depends on
vollib. Resolved: vollib in Table E,
py_vollib moved to the deprecated list. The rename
direction is consistent across every mention; at least one of the two
version numbers is wrong and neither is load-bearing.pmdarima — abandoned vs. slowing.
MiniMax's cluster summary enumerates it among avoid-list packages, but
its own tables mark it "slowing," it passes recency (2025-11-17), and it
appears in no avoid table. Resolved: slowing, retained in Table
E with the directive to prefer
statsforecast.AutoARIMA for new code.nautilus-trader — recommended vs.
unnecessary. MiniMax places it in reference-architecture layer
10 as a recommended engine and in its "not in primary stack"
table as something a USD 100 backtest does not need. Resolved:
retained at the execution-decision stage, matching Gemini's
assignment, with the honest note that it is optional at this scale and
carries the steepest learning curve in the survey.Free data is adequate for exploratory ingestion. It is not adequate
for a defensible point-in-time backtest, and the gap between those two
statements is where most retail quantitative work quietly dies.
[T6] (Source: MiniMax)
The governing distinction in this section is not whether a request
succeeds. It is whether the resulting dataset may lawfully be automated,
retained, redistributed, or used to sell a derived signal — and,
separately, whether the dataset reconstructs what was knowable at a
historical timestamp. [T6] (Source: MiniMax) Those
are three independent questions, and a source can pass the first while
failing both others. Yahoo Finance answers every HTTP request you send
it; that fact establishes nothing about your right to build a product on
the response.
Two definitions constrain every negative entry in Table F, and they are stricter than the definitions vendors use in marketing copy:
[T5] (Source:
MiniMax)[T5] (Source:
MiniMax)Where documentation does not establish a property, Table F records
unknown or not established, never
yes. A "no" in the point-in-time or survivorship column is a
data-engineering warning, not an accusation of inaccuracy: a provider
can report entirely correct current values while remaining unusable for
historical information-set reconstruction. [T6]
(Source: MiniMax)
Thirty-three distinct sources, merged across the surveyed reports.
PIT = point-in-time. SBF =
survivorship-bias-free.
| Source | Asset classes | History depth | Update latency | Rate limit | Auth required | PIT | SBF | TOS restriction on automated / derived use | Genuinely free |
|---|---|---|---|---|---|---|---|---|---|
SEC EDGAR (data.sec.gov) |
US filings, XBRL company facts, submissions, full-text search | Filings ~1990s onward; XBRL coverage begins later and varies by filer | Real time on filing acceptance | 10 req/sec documented fair-access ceiling | No key; descriptive User-Agent required | Partial — reconstructable if indexed by filing acceptance timestamp; the raw facts API is not itself a PIT database (amendments and taxonomy changes require event-time filtering) | Partial — filings of delisted issuers are retained, but EDGAR supplies no priced security master | Fair-access guidance restricts excessive automated requests; non-compliant clients are throttled or blocked. No redistribution restriction on the data itself | Yes |
FRED (api.stlouisfed.org) |
US and international macro, rates, labor, prices | Series-specific; longest series extend into the early 20th century | On release; revisions arrive after release | 120 req/min | API key | No on default endpoints; yes only
when realtime_start / realtime_end /
vintage_dates are used |
n/a (macro) | Terms permit broad public use; third-party series retain source restrictions; use the official API rather than scraping | Yes |
ALFRED (alfred.stlouisfed.org) |
Archived vintages of FRED macro and rates series | Series-dependent; often decades of vintages | Vintage snapshots at release | Same as FRED | API key | Yes for series with archived vintages | n/a | Same FRED terms and source-series restrictions | Yes |
US Treasury (fiscaldata.treasury.gov,
home.treasury.gov) |
Par yield curves, bill rates, auctions, debt, receipts and outlays | Dataset-specific, often decades | Daily, monthly, or event-driven | No universal documented quota; paginate and cache | None for public endpoints | No for revised series; record timestamps vary by dataset | n/a | US government data broadly reusable; third-party marks and dataset notices apply | Yes |
BLS (bls.gov/developers) |
CPI, employment, wages, productivity, release calendar | Series-specific, usually decades | Scheduled release; revision policy varies | Documented daily request and row limits — batch and cache | None for v1; registration raises limits | No unless release vintages are stored | n/a | Reusable with attribution; preserve release metadata | Yes |
BEA (apps.bea.gov/API) |
GDP, personal income, trade, industry accounts | National-accounts series, often decades | Scheduled releases with revisions | Quota tied to the API account | API key | No without vintage capture | n/a | Reusable; third-party inputs and trademarks may differ | Yes |
ECB Data Portal
(data.ecb.europa.eu) |
Euro-area rates, FX, macro, banking and financial statistics | Dataset-specific, frequently decades | Release and event dependent; revisions occur | No single documented universal quota | Usually none for public downloads | No unless vintage/release metadata is retained | n/a | ECB legal notices and dataset-specific reuse terms apply | Yes |
Bank of England IADB
(bankofengland.co.uk/boeapps/database) |
UK rates, yield curves, macro and financial series | Dataset-specific, often decades | Daily or release-based | None verified | Usually none | No without vintage storage | n/a | Bank terms and copyright notices apply | Yes |
Yahoo Finance / yfinance
(query1/query2.finance.yahoo.com) |
US and global equities, ETFs, FX, crypto, options chains, fundamentals, news | ~30y daily, ~60d intraday; provider guarantees no retention | Quotes near-real-time to 15-min delayed; historical-endpoint latency undocumented | No published limit; ~2,000 req/hr observed unofficially | None; cookie/crumb handshake varies and breaks | No | No | Yes — restricted. Terms contemplate personal/non-commercial use; automated extraction and redistribution restricted; interface is unofficial; IP bans reported | Freemium |
| Alpha Vantage | Global equities/ETFs, splits and dividends, fundamentals, earnings calendar and estimates, options, news/sentiment, FX, crypto, commodities, macro | 20+ years daily — but full daily history is a premium entitlement; the free tier is capped | Historical; delayed and real-time reserved to premium | Free quota is plan-dependent and revised without notice | API key | No | No | Yes. Terms and exchange-data policy restrict redistribution and commercial derived use | Freemium |
| Polygon.io | US equities/ETFs, options, futures, FX, crypto, corporate actions | Plan- and asset-dependent; the free tier is not a complete archive | Free tier delayed; latency is plan-dependent | Plan-specific | API key | No | No | Yes. Market-data licensing and plan terms restrict redistribution and commercial derived products | Freemium |
| Nasdaq Data Link (ex-Quandl) | Macro, fundamentals, equities, futures, rates — dataset-specific | Dataset-specific; many free datasets have finite history | End-of-day or periodic | Dataset-specific | API key for most datasets | No | No | Yes. Per-dataset license plus platform terms govern automated and derived commercial use | Freemium |
| Tiingo | Equities/ETFs, fundamentals, news, crypto | Plan-dependent; free account limited | Delayed or end-of-day by feed | Plan-dependent; no universal free limit | Token | No | No | Yes. Terms plus exchange redistribution restrictions; a free token does not imply commercial rights | Freemium |
| IEX Cloud | US equities, fundamentals, corporate actions | Unresolved | Unresolved | Unresolved | Token | No | No | Yes | Unresolved — do not adopt without verifying the product still exists |
| Marketstack | Global equities, EOD and intraday, corporate actions by plan | Free plan limited to recent history and request volume | Free tier delayed; plan-specific | Free-plan request quota is pricing-dependent | API access key | No | No | Yes. Commercial and redistribution rights are plan-dependent | Freemium |
| EOD Historical Data (EODHD) | Global equities/ETFs, corporate actions, fundamentals, calendars, options by plan | Plan-dependent; free and demo access limited | EOD or delayed; plan-dependent | Plan-specific | API token | No | No | Yes. Terms and dataset entitlements restrict automated redistribution and commercial derived use | Freemium |
| Stooq | Equities, indices, FX, futures, ETFs — daily history | Broad daily archives, instrument-dependent; no security master | End-of-day / delayed | No official API SLA or published limit located | Usually none | No | No | Yes. Automated-download permission must be checked before polling or redistributing; the access pattern is a scrape | Free for limited use — terms-sensitive |
| Financial Modeling Prep (FMP) | Fundamentals, analyst estimates, earnings calendars | Not established by any surveyed report | Not established | Not established | API key | No | No | Not established | Freemium (not established) |
| CoinGecko | Crypto prices, markets, exchanges, metadata | Endpoint- and plan-dependent; free history limited versus paid | Free public API delayed or rate-limited | Plan-specific; do not assume a permanent free number | Demo/public key depending on endpoint | No | No | Yes. Terms distinguish personal/free from commercial use and redistribution | Freemium |
| CryptoCompare | Crypto spot, OHLCV, trades, news, some derivatives | Asset-, exchange- and endpoint-dependent | Near-real-time or delayed by endpoint | Account- and endpoint-dependent | API key | No | No | Yes. Terms plus exchange-source rights restrict redistribution and commercial use | Freemium |
| Kaiko | Institutional crypto spot, derivatives, order books | Not free for production; trial or demo may be offered | Tick and order-book latency depends on paid plan | Contract-specific | Credential | Unknown | Unknown | Commercial license required | No |
Deribit (docs.deribit.com) |
BTC/ETH options, futures, order books, trades, instrument metadata | Exchange- and endpoint-dependent | Real-time REST and WebSocket | Exchange-specific published limits — check before polling | None for public market data; auth for private endpoints | No | No | Yes. Exchange terms control automated access and redistribution; public does not mean unrestricted commercial reuse | Free for public endpoints — terms-sensitive |
Binance (api.binance.com) |
Crypto spot and futures | 2017–present | Real-time REST and WebSocket | 1,200 req/min | Optional for public endpoints | Partial — printed trades are point-in-time by construction, but coverage is venue-scoped | No — venue-scoped; delisted pairs are not reliably retained | Automated access explicitly supported; redistribution governed by exchange terms. Geofenced for US IPs — Binance.US required | Yes for public endpoints |
Kraken (api.kraken.com) |
Crypto spot | Inception–present | Real-time REST and WebSocket | 1 req/sec public; REST OHLC returns max 720 bars per request | Optional for public endpoints | Partial — same venue-scoped caveat | No — venue-scoped | Automated access explicitly supported; redistribution governed by exchange terms | Yes for public endpoints |
Kalshi (docs.kalshi.com) |
CFTC-regulated event contracts — economic, political, climate, company; order books, trades, settlements | 2021–present; market-history retention is endpoint-specific with no blanket guarantee | Real-time REST and WebSocket; settlement after official resolution | Basic tier: token bucket, 200 read + 100 write tokens/sec,
default request cost 10 tokens ⇒ ~20 read / ~10 write req/sec. 429
responses omit Retry-After |
Account signup; RSA key for authenticated endpoints | No — archive locally | No | Automated access supported via a documented API; redistribution and derived commercial products restricted by terms; account eligibility applies | Yes (trading fees separate) |
| ForecastEx (via IBKR) | CFTC-regulated event contracts | Inception–present; no free historical L2/L3 order-book depth | Real-time via TWS/Gateway | 50 req/sec | IBKR account auth plus a locally running TWS/Gateway process | No | No | Permitted via the IBKR API; IBKR market-data terms apply | Yes with a funded IBKR account |
| Polymarket | Prediction-market prices, trades, order books, settlements | Market-specific; no documented universal archive guarantee | Near-real-time; endpoint behavior changes without notice | No stable published limit verified | Public endpoints may not require auth | No | No | Yes. Terms impose geographic, account, automated-use and IP restrictions; US participation requires separate legal verification | Public data free — access terms-sensitive and jurisdictionally gated |
GDELT (gdeltproject.org) |
Global news events, tone/sentiment, entity extraction | Multi-decade event corpus | Near-real-time updates | Not established | None | No | No | Openly accessible; source-specific notices apply. Not a licensed article-text feed | Yes |
| Econoday / Trading Economics | Economic-release calendars with consensus and actuals | Historical calendar depth is a paid, plan-gated feature | Event-time updates | Plan-specific for API access; free web access is not an API license | API key for the API | No | n/a | Yes. Automated scraping and commercial reuse restricted by provider terms | Freemium |
| Investing.com | Release calendars, prices, news | No authoritative archive guarantee | Web updates | No official public API | None | No | No | Yes. Automated scraping and redistribution restricted; unofficial clients break | Free web content — not a stable free API |
| CRSP | US equity prices, delisting returns, historical index constituents | Not stated by any surveyed report | n/a | n/a | Institutional subscription | Yes — the reference standard | Yes — the reference standard | Academic/institutional license; redistribution prohibited | No |
| Compustat / WRDS | Fundamentals, including point-in-time fundamentals products | Not stated by any surveyed report | n/a | n/a | Institutional subscription | Yes | Yes | Academic/institutional license | No |
| NYSE TAQ | US trade and quote tick data | Not stated by any surveyed report | n/a | n/a | Subscription | Yes | n/a | Commercial license | No |
Three counts worth reading off the table directly.
Survivorship-bias-free = yes appears on exactly three rows, and
all three are paid (CRSP, Compustat/WRDS, TAQ).
Point-in-time = yes appears on one free row (ALFRED),
plus five conditional-on-vintage-storage rows (FRED, BLS, BEA, ECB, Bank
of England) and one partial row (SEC EDGAR). Every market-data vendor in
the table is a no on both. [T6] (Source:
MiniMax, extended with Gemini)
The last three rows are in a "free data sources" table on purpose.
CRSP, Compustat/WRDS and NYSE TAQ are not recommendations — they are the
benchmark the free universe is being measured against, and their absence
from the budget is a material limitation rather than a reason to
substitute Yahoo data silently. [T6] (Source:
MiniMax)
Equity and ETF price history, splits and corporate
actions. Alpha Vantage, Nasdaq Data Link, Tiingo, EODHD,
Marketstack, Polygon, Stooq and Yahoo Finance cover overlapping
end-of-day universes. [T5] (Source: MiniMax) Alpha
Vantage documents raw and adjusted daily, weekly and monthly series with
split and dividend fields, and its daily history reaches 20+ years —
but full daily history and most intraday access are
premium, so the free key supports exploratory pulls, not broad
universe ingestion. [T5] (Source: MiniMax) The
defect set is identical across all of them: ticker changes, inconsistent
exchange identifiers, missing delisted instruments, changing adjustment
conventions, stale fundamentals, and silent vendor backfills.
[T5] (Source: MiniMax) No source surveyed
establishes a free, complete delisted-security archive, which
is why the survivorship column reads no for every vendor row.
[T6] (Source: MiniMax)
Fundamentals, earnings calendars and estimates. SEC
EDGAR is the authoritative free source, and it is the only one whose
structure supports point-in-time reconstruction: accession numbers,
filing dates, amendment filings and company submissions are all
preserved, so an ingestion system that stores the original filing and
its publication timestamp can rebuild the historical information set.
[T5] (Source: MiniMax, Gemini) The raw facts API
is not itself point-in-time — later amendments and taxonomy changes
require event-time filtering, and XBRL tagging is inconsistent across
filers. [T5] (Source: MiniMax, Gemini) Alpha
Vantage, Nasdaq Data Link, Tiingo, EODHD and Financial Modeling Prep
expose vendor-normalized fundamentals, estimates and earnings calendars,
but normalization and restatement policy are not the same thing as
point-in-time data. [T5] (Source: MiniMax) Analyst
estimates are the most exposed series in the entire survey to
historical-survivor and revision bias, and no source surveyed
establishes a complete historical consensus-vintage archive.
[T6] (Source: MiniMax)
Options chains and implied volatility — the largest gap in
the free universe. Alpha Vantage documents real-time and
historical US options, put-call ratios and volume/open-interest ratios,
and marks all of them premium. [T5] (Source:
MiniMax) Yahoo Finance exposes options through an unofficial
interface with no stable contract and no verified free historical-chain
guarantee. [T6] (Source: MiniMax) Deribit
publishes public crypto options instruments, trades, order books and
ticker data, but it is exchange-specific and not a US equity-options
substitute. [T5] (Source: MiniMax) Gemini's survey
names no options source at all. No source surveyed provides a
free, documented, long-history, point-in-time US options chain and
volatility surface. [T6] (Source:
MiniMax)
The methodological point underneath that gap matters more than the
gap itself: a chain snapshot is not a volatility
history. A usable IV surface requires timestamped bid/ask,
contract metadata, corporate-action-adjusted strikes, open interest,
settlements, and the delisted and expired contracts. Free sources
generally supply a current chain or a short rolling window, which is a
different object entirely. [T6] (Source:
MiniMax)
Macro, Treasury, rates and central-bank data. FRED
and ALFRED together give broad series access with genuinely
vintage-oriented workflows; Treasury publishes daily par yields, bill
rates and related datasets; BLS and BEA publish release-based economic
data; the ECB and the Bank of England publish euro-area and UK datasets.
[T5] (Source: MiniMax, Gemini) These are
the strongest genuinely free sources in the entire survey.
[T6] (Source: MiniMax) Their defining defect is
revision: the current observation is not the observation that was
available on the historical decision date. Any pipeline that reads
FRED's default endpoints and stores one value per series has already
destroyed its own point-in-time property. Storing
release_timestamp, observation_date,
vintage_date and source_series_id is the
entire remedy, and FRED's vintage_dates parameter — the
same facility ALFRED exposes — is how you get it. [T5]
(Source: MiniMax, Gemini)
SEC EDGAR access mechanics. EDGAR requires a
declared, descriptive User-Agent and a reasonable request rate under SEC
fair-access guidance; the documented ceiling is 10 requests per second,
and automated clients that ignore the guidance are throttled or blocked.
[T5] (Source: Gemini, MiniMax) This is not
theoretical. MiniMax reports receiving a 403 from the SEC during its own
research pass — operational evidence that the access controls are live,
not that the service is unavailable to compliant clients.
[T5] (Source: MiniMax)
News and sentiment. Alpha Vantage documents a
news-and-sentiment endpoint whose availability is entitlement-dependent;
Nasdaq Data Link, Tiingo and CryptoCompare expose news datasets with
variable historical depth and redistribution rights; GDELT provides
openly accessible global news-event and tone data. [T5]
(Source: MiniMax) GDELT is an event-extraction corpus, not a
clean licensed financial-news feed, and its known defects — source
duplication, language imbalance, timestamp ambiguity, entity-resolution
errors — require validation before any signal is derived from it.
[T5] (Source: MiniMax) No source surveyed
establishes complete historical news survivorship, stable article-text
licensing, or bias-free sentiment labels. Treat every sentiment
field as a vendor-derived feature, never as ground truth.
[T6] (Source: MiniMax) Gemini's survey names no
news source at all.
Crypto market data. CoinGecko and CryptoCompare are
the practical free and freemium aggregators for spot prices, exchange
metadata, OHLCV and market-cap fields; their aggregation defects are
heterogeneous clocks, venue outages, symbol-mapping drift and
wash-trading exposure, and neither is point-in-time or
survivorship-bias-free by default. [T5] (Source:
MiniMax) Exchange-native access is cleaner and narrower: Binance
offers 1,200 req/min across spot and futures from 2017 but geofences US
IPs, Kraken offers 1 req/sec public with a 720-bar cap per REST OHLC
call, and Deribit covers its own BTC/ETH derivatives book.
[T5] (Source: Gemini, MiniMax) Exchange feeds
cover only their own venue and their own listed-instrument lifecycle,
which is precisely why they cannot be survivorship-bias-free at the
asset-universe level. [T6] (Source: MiniMax, applied to
Gemini's rows) Kaiko is the institutional alternative and it is not
free. [T5] (Source: MiniMax)
Prediction markets and event contracts. Kalshi is
the only venue in this class with a real, documented API contract:
public market data, order books, trades, and market status and history
surfaces, with authenticated trading. [T5] (Source:
MiniMax, Gemini) Its Basic tier runs a token bucket rather than a
fixed window — 200 read tokens/sec and 100 write tokens/sec against a
default request cost of 10 tokens, which works out to roughly 20 read
and 10 write requests per second — and its 429 responses currently omit
retry metadata, so client backoff must be self-managed.
[T5] (Source: MiniMax, corroborated by Gemini)
ForecastEx is reachable at 50 req/sec but only through a locally running
IBKR TWS or Gateway process, which makes it an infrastructure dependency
rather than a plain HTTP endpoint. [T5] (Source:
Gemini) Polymarket publishes public market information, but its
terms impose geographic, account and automated-use restrictions, its
public interfaces change, and no surveyed report verified a compliant
US-retail authorization from the public documentation. Do not infer
legality or unrestricted automated rights from the existence of public
JSON. [T6] (Source: MiniMax)
For every venue in this class, historical settlement and
order-book completeness is endpoint-specific with no blanket retention
guarantee. The operational consequence is unambiguous: archive
every market, series, close time, settlement value and rule text
locally, at capture time, because contract definitions and market
availability do not constitute a stable historical database.
[T6] (Source: MiniMax)
Economic-release calendars. The genuinely free
primary sources are the BLS release calendar, BEA release schedules,
Federal Reserve calendars, Treasury auction schedules and central-bank
calendars — all of which give event timing and none of which give
historical consensus forecasts. [T5] (Source:
MiniMax) The richer consensus-and-actual fields sit behind Trading
Economics, Econoday, Investing.com and broker calendars, all plan-gated
or terms-sensitive. [T5] (Source: MiniMax) The
constraint that matters: a calendar without archived "what was
known when" consensus cannot support an event-surprise
backtest. Knowing that CPI printed on a date tells you nothing
about the surprise unless you also stored the forecast that existed
before the print. [T6] (Source: MiniMax)
Reproduced as an explicit list, because this is the set with
downstream legal exposure. [T5] (Source:
MiniMax)
tos_restricts_automation: true.The most defensible free inputs, by contrast, are the government and
public sources: FRED, ALFRED, Treasury, BLS, BEA, SEC, ECB, Bank
of England and GDELT — subject to source-specific notices,
fair-access controls and revision handling, but not to redistribution
licensing. [T5] (Source: MiniMax)
One caveat on the whole list: no surveyed report quotes or links a
specific terms clause for any of these restrictions. Every TOS cell in
Table F is an unlinked paraphrase of a landing page. Treat the list as a
legal-review worklist, not as legal advice. [T6]
(Source: MiniMax digest, citation-integrity flag G)
Interfaces classified as unstable or unofficial, meaning an ingestion
pipeline built on them should carry an identified replacement feed from
day one. [T6] (Source: MiniMax, Gemini)
Where no official stability statement was found, this classification
is a statement of operational risk, not a prediction of imminent
failure. [T6] (Source: MiniMax)
Six requirements follow directly, and none of them is optional if the
output is meant to survive scrutiny. [T6] (Source:
MiniMax)
observation_date,
release_timestamp, retrieval_timestamp,
vintage_date and effective_timestamp, and
never overwrite a prior macro observation or filing
representation.current / delayed /
historical / vintage / adjusted /
raw. Adjusted OHLCV is not corporate-action history.Only two of the three surveyed reports addressed this domain at all.
Qwen's report contains zero coverage — no Table F, no
named provider, no endpoint, across roughly 6,700 words, despite
proposing a system that "ingests real-time odds from a prediction
market." Every claim in this section therefore rests on a two-source
base, and the majority-rule tiebreak available elsewhere in this report
is unavailable here. [T6] (Source: Qwen digest,
§Cluster 8)
Neither surviving report verified every current commercial term, plan
quota, retention policy and endpoint-specific archive; MiniMax says so
explicitly, and its self-reported 403 from the SEC means an unknown
subset of its cells were written from prior knowledge rather than a live
fetch. Every "no published limit" cell in Table F inherits that
uncertainty. [T6] (Source: MiniMax digest,
citation-integrity flag H) Run a terms review and a live
schema-and-limit probe before any of this reaches production.
The surveyed reports also leave named-source gaps that this section
does not fill, because filling them from outside the source base would
be fabrication: Finnhub, Twelve Data, Databento, Alpaca, OpenBB, CBOE's
own free index and VIX term-structure data, OCC/LiveVol, Coinbase's
public REST API, the Federal Reserve's H.15 and Data Download Program as
distinct from FRED, IMF, World Bank, OECD, Eurostat, exchange
corporate-action feeds, CFTC Commitments of Traders, and the
non-Kalshi/non-Polymarket prediction venues. The CBOE omission is the
most conspicuous, since both reports declare options data the largest
free-data gap while neither examines the exchange that publishes free
volatility indices. [T6] (Source: MiniMax digest, flag
K)
Tiebreak rule applied throughout: where MiniMax and Gemini disagree on a data-source attribute, MiniMax governs. MiniMax's cluster-8 digest scores 8 / 8 / 7 on completeness, correctness and thoroughness and applies explicit definitions of point-in-time and survivorship-bias-free before assigning any value; Gemini's own digest identifies cluster 8 as its weakest section (7 sources, two required domains missing) and flags its quantitative claims as presented without stated method or citation. Qwen contributes nothing to this section and cannot break ties.
vintage_dates). MiniMax: no unless vintage dates are
stored; lists ALFRED separately as the PIT = yes row.
Reconciled: these describe the same facility. FRED's
default endpoints are not point-in-time; the vintage_dates
/ realtime_start / realtime_end parameters and
ALFRED are the same vintage interface. The qualifier is preserved inline
in Table F because dropping it would license exactly the error the
column exists to prevent.Cross-source consistency adjustment (not a
conflict). Gemini marks Binance and Kraken point-in-time = YES
and survivorship-bias-free = YES. MiniMax never evaluated either venue,
so there is no disagreement to adjudicate. Both rows are nonetheless
downgraded to PIT = partial and SBF = no by applying MiniMax's stated
principle that exchange-native sources cover only their own venue and
their own listed-instrument lifecycle — delisted trading pairs are not
reliably retained on exchange endpoints, which defeats
survivorship-freedom at the asset-universe level. This is synthesis
across the two reports, tagged [T6], and is logged
separately so the audit trail does not overstate what the sources
said.
Scope and standing caution. This section merges three independently commissioned research reports (Qwen, Gemini, MiniMax) against U.S. federal law and the law of the Commonwealth of Massachusetts, at a stated research date of August 1, 2026. It is a synthesis of secondary research, not legal or tax advice, and it is not a substitute for a licensed practitioner's opinion. Two structural facts govern how much weight any single line here can bear. First, all three source reports exhibit citation defects at load-bearing points, and one of the three (Qwen) omits every tax authority entirely. Second, and more seriously, MiniMax's own retrieval log records that no Massachusetts primary source was successfully retrieved — mass.gov returned access errors for the Securities Division, the Gaming Commission, the Attorney General, and 830 CMR — so every Massachusetts statutory citation reaching this section arrives unverified against the primary text (Source: MiniMax). Where the three reports conflict on a legal or tax question, this section applies a deliberate asymmetry: it prefers the more specifically cited claim over the more confidently stated one, and it defaults to unsettled unless two sources independently pin the same authority. Claims that could not be pinned to a specific rule, statute, or form number were downgraded or dropped outright; those decisions are itemized in
### Resolved Conflictsat the end.
Evidence tier key as used below: [T1]
peer-reviewed and replicated; [T2] peer-reviewed, single
study; [T3] working paper, preprint, or agency technical
release; [T4] grey literature, trade press, industry
convention; [T5] primary source — statute, regulation, SRO
rule, court order, IRS form or publication; [T6] author
inference, unverified, or genuinely unknown. Per the brief, every legal
or regulatory claim below carries [T5]
only where a specific rule, statute, or form number
survived conflict resolution. Claims that cannot be pinned to a specific
citation carry [T6] and are stated as unsettled.
The SEC does not register the investor. It registers the intermediary, and the practical consequence for a USD 100 self-directed account is that every SEC obligation in this subsection falls on the broker-dealer, not on the principal.
Broker-dealer registration. The Securities Exchange
Act of 1934 § 15, codified at 15 U.S.C. § 78o, requires
broker-dealers to register with the Commission [T5]
(Source: MiniMax; retrieved Aug 1, 2026). A FINRA-member,
SEC-registered broker-dealer is therefore the necessary counterparty to
any equity, ETF, or listed-option position in this experiment. The
investor's own registration obligation is nil.
Regulation Best Interest. Reg BI, at 17 CFR
§ 240.15l-1, obligates a broker-dealer to act in a retail
customer's best interest when making a recommendation [T5]
(Source: Gemini; retrieved Aug 1, 2026). Only Gemini cites this
rule; MiniMax does not raise it. Treat it as accurate but
single-sourced. Its operative consequence at USD 100 is negligible — Reg
BI attaches to recommendations, and a purely self-directed
account receives none. It does not bar a cash account and imposes no
capital gate [T5].
Settlement. The amendment to SEC Rule
15c6-1 moved standard U.S. securities settlement to
T+1, effective May 28, 2024
[T5] (Source: MiniMax, Gemini). Qwen asserts T+2
in three separate places and builds its equity-turnover analysis on it;
that is stale by roughly two years and is corrected here. The correction
is not cosmetic — T+1 approximately doubles the achievable round-trip
frequency in a cash account relative to Qwen's model.
Jurisdictional boundary — where the SEC does not
reach. MiniMax states affirmatively that event contracts,
futures, spot crypto, and crypto derivatives traded on CFTC-registered
venues are not directly SEC-jurisdictional, while
equity ETFs, listed options, security-based swaps, and security futures
are [T5] (Source: MiniMax). Qwen
adds that the SEC/CFTC split over crypto assets turns on whether the
asset satisfies the Howey test — securities to the SEC,
commodities and their derivatives to the CFTC — citing a March 2026
joint SEC/CFTC crypto-asset guidance summary [T4]
(Source: Qwen). The Howey framing is standard and uncontested
across the sources; the specific 2026 guidance document reaches this
section only through a law-firm client alert, so the boundary principle
is [T5] but the 2026 guidance itself is
[T4].
Investment Advisers Act. Covered in §11.9 below, since the analysis is jointly federal and Massachusetts.
Know Your Customer and the Customer Identification
Program. FINRA Rule 2090 (Know Your Customer)
obligates member firms to use reasonable diligence to know the essential
facts of every customer [T5] (Source: MiniMax).
Account opening in either a cash or margin account requires SSN or ITIN,
current address, employment, a financial profile, and a risk-tolerance
disclosure [T5] (Source: MiniMax; retrieved Aug 1,
2026).
The CIP citation is contested and is not resolved
here. MiniMax cites 17 CFR § 1010.230; Gemini
cites 31 CFR § 1020.220 (Sources: MiniMax,
Gemini). These cannot both be right, and MiniMax's is wrong on its
face — Bank Secrecy Act CIP rules live in 31 CFR,
administered by FinCEN, not in 17 CFR. Gemini's Part 1020 is the
banks subpart, and the broker-dealer CIP subpart is Part 1023.
No source independently cites the broker-dealer subpart directly.
The exact CIP section number is therefore marked
[T6] and left unresolved — the
requirement (a written CIP, identity verification at account
opening) is settled and unanimous, but the section number that would let
a reader look it up is not established by any of the three reports. Both
sources invoke the USA PATRIOT Act with no section number; that
reference is dropped.
Suitability. FINRA Rule 2111
imposes a firm-level obligation to determine that a recommendation is
suitable to the customer's investment profile [T5]
(Source: MiniMax, Gemini). Rule 3110 governs
supervision [T5] (Source: MiniMax). The operative
point at USD 100: a purely self-directed, execution-only account
receives no recommendation and therefore does not trigger the
suitability obligation at all [T5]. MiniMax attributes this
carve-out to a specific supplementary paragraph,
FINRA Rule 2111.05; that pincite is not corroborated and
the substantive conclusion flows more cleanly from the absence of a
recommendation than from an express exclusion, so the .05
pincite is downgraded to [T6] while the conclusion stands
at [T5]. MiniMax's enumeration also lists
FINRA Rule 2110 in a mandatory table with no proposition
attached anywhere in its report; that citation is
dropped as unsupported.
The material exception: option and margin approvals
inherently involve firm-level suitability review, because the
firm must affirmatively approve the account for those privileges
[T5] (Source: MiniMax). Self-direction does not
bypass that gate.
Options approval — the tier ladder is convention, not
regulation. This is a resolved conflict worth stating plainly.
MiniMax presents Levels 1–4 as "codified" at
FINRA Rule 2360(b)(11)–(12); Qwen refers to "Tier 1" and
"Tier 4" attributed to broker internal policy. Rule 2360 governs
options account approval and firm diligence [T5],
but the four-level ladder itself is an industry convention set firm by
firm, not a FINRA-codified schedule [T4] (Sources:
MiniMax, Qwen; resolved toward Qwen's "internal policy"
characterization). MiniMax separately cites
FINRA Rule 2360(b)(16) for uncovered short option writing
[T5]. Presented as convention, the ladder runs:
| Level | Privileges (industry convention, not a FINRA schedule) | Realistic at USD 100? |
|---|---|---|
| 1 | Covered call writing; cash-secured put writing | Grantable, but requires stock or cash collateral the account does not have |
| 2 | Buying calls and puts, plus Level 1 | Yes — this is the operative level |
| 3 | Spreads, uncovered writing, married puts; many firms require a margin account | No (margin infeasible — see below) |
| 4 | Uncovered/naked writing; strict Reg T or portfolio-margin requirements | No |
The two sources conflict on whether Level 3 is
reachable: MiniMax's §A.2 states a USD 100 account "cannot
reasonably reach Level 3+," while its own §B.2 states Level 3 at
Interactive Brokers is "approvable immediately"; Qwen states spreads
"are unlikely to be approved for an account with only $100 in equity."
Two of three positions (Qwen plus MiniMax's §A.2) agree, and the
constraint is arithmetic rather than discretionary — most firms require
a margin account for spreads, and margin has a USD 2,000 floor.
Resolution: Level 2 is the practical ceiling at USD 100
[T5] for the margin arithmetic, [T4] for the
approval convention.
Minimum viable option position. The sources disagree
by an order of magnitude — Qwen gives "several dollars to over $100" per
contract; MiniMax gives both "USD 50–200 per contract" and, three
sections later, "USD 10 per contract minimum." No source supplies an
underlying, strike, expiry, or quote date for any of these figures.
All per-contract cost figures are marked
[T6]. MiniMax's claim that micro-options render
the options path feasible at USD 100 is self-flagged [T6]
in its own text and remains [T6] here.
Pattern day trader rule — citation resolved, current status
not. Qwen cites "SEC Rule 2222" as the PDT authority.
No such rule exists; this is a fabricated citation in a
mandatory regulatory table and is dropped. Qwen also
inverts the rule's logic, describing it as applying "only to margin
accounts with equity above $25,000" when the rule in fact restricts
accounts below that threshold. Gemini supplies the correct
citation: FINRA Rule 4210(f)(8)(B) (with Rule 2520 as
the NYSE-legacy analogue), requiring a pattern day trader — four or more
day trades in five business days — to maintain USD 25,000 minimum equity
[T5] (Source: Gemini; resolved against Qwen on citation
specificity).
MiniMax asserts the PDT rule was rescinded effective June 4,
2026, replaced by intraday margin standards under an April 2026
SEC approval, with phase-in through October 2027. That claim reaches the
section through Wikipedia citing "FINRA Notice 26-10" —
a tertiary chain, with no primary notice retrieved, and MiniMax's own
text concedes some brokers may still enforce the USD 25,000 floor
pending system updates. Per the cautious-default rule, the
rescission is marked [T6]/unsettled, and this
section does not rely on it. It does not matter operationally: the PDT
rule attaches to margin accounts, and margin is unavailable at
USD 100 regardless of which regime is in force.
Margin is infeasible, and this is the hardest constraint in
the section. Regulation T (12 CFR Part 220)
sets initial margin at 50% of the purchase price of marginable equity
securities, and FINRA Rule 4210(b)(4) sets the minimum
equity to open a margin account at USD 2,000
[T5] (Source: MiniMax). A USD 100 account is 5% of
the way to the floor. Gemini reaches the identical conclusion from the
PDT side: the account must operate strictly as a cash
account [T5] (Source: Gemini). This is the one
point in the entire section where all sources that address it agree
without qualification.
Cash-account mechanics and the free-riding trap. In
a cash account every transaction must be fully paid for. Selling a
security before the purchase that acquired it has settled is a
free-riding violation; the penalty is a 90-day
restriction under which the account may only purchase with
settled cash [T5] (Source: MiniMax, Qwen —
independently stated). MiniMax attributes the three-strikes
threshold — three good-faith violations in twelve months triggering the
freeze — to SEC Office of Investor Education material and Reg T
[T4]; the existence of the 90-day restriction is
corroborated by two sources and stands at [T5], but
the "three violations in twelve months" count is single-sourced
to investor-education material rather than rule text and is marked
[T4].
Practical consequence, corrected for T+1: the account can execute roughly one round trip every two business days without touching unsettled proceeds, not one every three days (MiniMax, which assumed a longer cycle) and not one every two days because of T+2 (Qwen, which had the settlement cycle wrong and reached a similar number by coincidence).
Broker reporting cadence. The broker-generated
Form 1099-B already reflects wash-sale loss disallowed
in Box 1g for covered securities [T5]
(Source: MiniMax; 2026 Instructions, retrieved Aug 1, 2026).
Spot crypto is reported on Form 1099-DA for brokers in
scope beginning with 2025 transactions [T5] (Source:
MiniMax). CFDs and off-broker spot crypto are generally not 1099-B
reportable [T5].
Account-opening latency. MiniMax gives 1–5 business
days typical, with manual review extending to 1–4 weeks, and notes that
major retail brokers offer instant-to-one-day automated onboarding as of
the research date [T4] (Source: MiniMax). Qwen
makes the same point qualitatively, flagging KYC/AML as a structural
gate that consumes time off the 90-day clock [T4]
(Source: Qwen). Both are broker-operational observations rather
than rule requirements, so both sit at [T4]. Budget the
clock accordingly: onboarding is not free of the 90 days.
Jurisdiction. The CFTC's exclusive jurisdiction over
futures contracts on commodities sits at CEA § 2(a)(1)(A), 7
U.S.C. § 2(a)(1)(A) [T5] (Source: Gemini;
corrected against MiniMax). MiniMax cites the same 7 U.S.C.
provision but labels it "Securities Act § 2(a)(1)(A)" — a Title 7
Commodity Exchange Act provision given a securities-statute name.
Gemini's labeling is correct and is adopted. Swaps fall under
CEA §§ 1a(47) and 2(h) [T5] (Source:
MiniMax).
Event contracts — the operative prohibition.
CEA § 5c(c)(5)(C) prohibits designated contract markets
from listing event contracts involving terrorism, assassination, war,
gaming, or unlawful activity, subject to Commission review; 17
CFR § 40.11 implements it, with § 40.11(a)(1) carrying the
prohibition and a review-and-approval mechanism running 90 days
[T5] (Source: MiniMax, Gemini — independently and
identically cited). This is one of the few citations in the entire
section that two sources pin to the same authority with the same
subsection, and it is accordingly the most reliable regulatory citation
here. MiniMax attributes the 90-day period to both the statute and the
regulation without resolving which; that sub-question is
[T6].
Pending rulemaking. MiniMax reports a CFTC notice of
proposed rulemaking at 91 FR 35806 (June 12, 2026),
release 9249-26, proposing a "Reg 40.11 Appendix F" framework for
evaluating whether an event contract involves an enumerated activity or
is contrary to the public interest. MiniMax states it is in
public-comment phase and not adopted [T3]
(Source: MiniMax). Neither other source mentions it. Its final
text and adoption are [T6]/unknown. A rule that is not
adopted binds no one, and this section does not rely on it — but a
reader planning a 90-day window should know that the framework governing
event-contract legality was actively in flux at the research date.
Litigation history — the central conflict of this section. The three reports give materially different accounts of whether a Massachusetts resident may lawfully trade CFTC-regulated event contracts. This is resolved in §11.5 below, because the operative question is a state-jurisdiction question. What belongs here is the federal record:
[T2]
(Source: Gemini). This is the most completely formed case
citation supplied by any of the three reports — court, docket, date, and
reporter — and on citation specificity it is the one to prefer. What it
establishes, however, is that the CFTC could not block
those contracts under its own review authority. Gemini extends it to
"Kalshi and ForecastEx are 100% lawful venues for MA residents," which
the cited holding does not support and which is contradicted by two
other sources and by a later Massachusetts court order. The
citation is retained; the extension to state-law preemption is
rejected — see §11.5.[T3] (Source: MiniMax). MiniMax's own digest flags
this release-number series as unverified and specifically notes one
release in the series (9276-26) appears in its bibliography supporting
nothing in its body — a hallucination signal. The entire 92xx-26
release-number series is marked [T6] pending primary
verification, including 9240-26, which anchors the BTCPERP
§1256 conclusion in §11.7.Venues. KalshiEX, LLC holds CFTC designation as a
contract market under CEA § 5, granted November
2020 [T4] (Source: MiniMax — sourced to
Wikipedia and the CFTC DCM list; the DCM list is primary but was not
retrieved). Polymarket acquired QCEX, a CFTC-licensed derivatives
exchange and clearinghouse, for a reported USD 112 million in 2025 and
received an Amended Order of Designation in November
2025, with a DOJ/CFTC probe closed without charges in July 2025
[T4] (Source: MiniMax). MiniMax supplies no order
number and no release number for the Polymarket designation despite
citing six numbered releases elsewhere, so Polymarket's federal
authorization — which MiniMax itself labels its most settled headline
finding — rests on the weakest citation in its report. Marked
[T4], not [T5]. ForecastEx, LLC is
described as a CFTC-registered DCM related to the Cantor Fitzgerald
group operating economic and financial event contracts, but
MiniMax concedes it did not retrieve the designation
order; Gemini names ForecastEx as lawful without citing its
designation at all. ForecastEx's registration status is
[T6]/unverified — and it should not be carrying
tax-treatment rows in any table on the strength of an unretrieved
designation.
Interactive Brokers launched a multi-venue
event-contracts platform in May 2026 aggregating
KalshiEX, CME Group, and ForecastEx contracts [T4]
(Source: MiniMax, via Wikipedia).
Bitcoin perpetual futures. MiniMax reports the CFTC
approved KalshiEX's BTCPERP contract, classified by the
Commission as a futures contract under CEA § 5c(c)(4)
and Reg 40.3 [T6] (Source: MiniMax). This
classification matters enormously — it is the bridge to Section 1256
treatment in §11.7 — and it rests on release 9240-26, which is in the
unverified series above. MiniMax separately cites
CEA § 5c(c)(4) as the source of mark-to-market for DCM
event contracts; CEA § 5c(c) concerns self-certification and
Commission stay procedures, not mark-to-market, so that cite is
dropped and the mark-to-market premise loses its only
stated support.
Crypto derivatives margin. MiniMax cites
CEA Reg 41.42–41.49 for margin on regulated crypto
derivatives; 17 CFR Part 41 concerns security futures
products customer margin, jointly administered with the SEC,
which is a different instrument class. The citation is
dropped [T6]. The operative fact survives on its
own: CME standard Bitcoin futures carry minimum notional in the USD
5,000–25,000 range, far exceeding a USD 100 stake [T4]
(Source: MiniMax). BTCPERP minimum notional is unknown
[T6].
Account age. MiniMax cites
Dodd-Frank § 745 for an 18+ age threshold at CFTC venues,
then immediately downgrades it in its own text to "per most operator
interpretations" and "conservative inference." Dodd-Frank § 745
concerns DCM core principles, not customer age. The citation is
dropped; the
federal-venue-versus-Massachusetts-gambling-age question (18 vs. 21) is
[T6]/unknown and no source resolves it (Source:
MiniMax).
Governing statute — conflict resolved. MiniMax cites
M.G.L. c. 110H as the Massachusetts Securities Act
across roughly fourteen load-bearing uses, and in the same passage also
cites "c. 110" and "c. 110, § 410," producing three chapter designations
in one report. Gemini cites M.G.L. c. 110A
[T5] (Sources: Gemini vs. MiniMax; resolved toward
Gemini). Two independent grounds favor Gemini: it is internally
consistent, and it is corroborated by a reported appellate decision
under the same chapter (below). Compounding the case against MiniMax,
its own mandatory table supplies federal URLs —
law.cornell.edu/uscode/text/15/78o and
.../80b-2 — as the source for its two Massachusetts rows,
which is an admission that no state authority was retrieved. c.
110A is adopted; c. 110H and c. 110 § 410 are dropped.
Section-level pincites within c. 110A from MiniMax (§§ 1, 2, 6, 11,
14–18) inherit the chapter defect and are not
propagated.
Massachusetts imposes a fiduciary duty above the FINRA
baseline. This is a direct conflict and it resolves against
MiniMax. MiniMax states the Division "does not impose a
separate suitability obligation beyond the FINRA Rule 2111 baseline"
[T6] — an unsourced negative. Gemini cites 950 CMR
12.207, the Massachusetts state fiduciary-duty regulation
imposing a duty of utmost care and loyalty on broker-dealers dealing
with Massachusetts retail customers, and cites Robinhood
Financial LLC v. Secretary of the Commonwealth, 492 Mass. 696
(2023) — a Supreme Judicial Court decision with a full
official-reporter citation — upholding it
[T5]/[T2] (Source: Gemini). A
specific regulation plus a reported SJC decision defeats an unsourced
negative. Massachusetts broker-dealers owe Massachusetts retail
customers a fiduciary standard, not merely FINRA suitability.
The practical consequence at USD 100 runs in the investor's favor and
is worth naming: 950 CMR 12.207 is a constraint on the broker,
and Gemini reads it as protective against gamification and against
overly permissive options approvals [T5]. It creates no
filing obligation for the principal. It may, however, make a
Massachusetts firm more conservative in granting option
privileges than the national convention in §11.2 suggests — which cuts
against reaching even Level 2 quickly.
No Massachusetts-specific net-worth minimum for retail
options. MiniMax states Massachusetts imposes no minimum net
worth for retail options trading different from the FINRA baseline
[T6] (Source: MiniMax). This is an uncorroborated
negative finding from the report whose chapter citation failed; it is
retained as [T6] rather than [T5] and should
not be relied on without checking 950 CMR 12.
No event-contract carve-out located. MiniMax states
that no Securities Division carve-out or guidance specific to event
contracts was located as of the research date, and that the Division's
role has been eclipsed by the CFTC/Attorney General dialogue
[T6] (Source: MiniMax). Gemini and Qwen are
silent. Absence of located guidance is not evidence of absence
of guidance, particularly where the same report concedes
mass.gov retrieval was blocked. [T6]/unknown.
Margin-rate cap. MiniMax cites
M.G.L. c. 167 § 15A for a 9.5% simple interest cap on small
margin loans, and appends its own flag: "primary verification needed."
Since margin is infeasible at USD 100 in any case, the claim is
inoperative here. Marked [T6] and not relied
upon (Source: MiniMax).
Statutory framework. The Massachusetts Gaming
Commission derives authority over casino and slots gaming from
M.G.L. c. 23K (Massachusetts Gaming Act, 2011)
[T5] (Source: MiniMax, Gemini). Sports wagering is
governed by M.G.L. c. 23N, with Gemini pinciting
§ 3 and characterizing it as regulating wagering on
athletic contests [T5] (Source: Gemini).
MiniMax gives c. 23N a different and incompatible
identity — the "Fantasy Contest Act, 2016," administered by a
"Massachusetts Fantasy Contest Commission" — and attributes 2022 sports
wagering instead to "M.G.L. c. 23O." Resolution:
Gemini. Gemini's is the more specific citation (chapter plus
section), it is internally consistent, and MiniMax's alternative
introduces a chapter (23O) that its own digest flags as non-existent and
a commission whose existence is uncorroborated. c. 23N § 3 is
adopted as the sports-wagering authority; c. 23O and the "Fantasy
Contest Commission" are dropped. Note the collateral damage:
MiniMax states the Superior Court in Commonwealth v. KalshiEX
relied on "c. 23K and c. 23N together," and if MiniMax's identification
of c. 23N was inverted, its characterization of the court's reasoning is
unreliable. That reasoning is therefore [T6].
MiniMax's Category 1/2/3 license ladder (resort casinos, slots-only,
sports wagering) conflates the c. 23K casino category scheme with the
separately numbered sports-wagering license categories, and its operator
list is self-flagged as unstable. Marked [T4] and
not load-bearing.
The jurisdictional boundary — the single most consequential unresolved question in this section.
The three reports give three incompatible answers to whether a Massachusetts resident may lawfully trade CFTC-regulated event contracts:
| Source | Position | Support offered |
|---|---|---|
| Gemini | CFTC DCM event contracts on inflation, CPI, interest rates, and elections are federally preempted under CEA § 2(a)(1)(A) and fall outside MGC jurisdiction. "Kalshi and ForecastEx are 100% lawful venues for MA residents." | KalshiEX LLC v. CFTC (D.D.C. 2024, aff'd D.C. Cir. 2024)
[T2]; CEA § 2(a)(1)(A) [T5] |
| MiniMax | Contested/unsettled. MGC "has not asserted direct jurisdiction" over CFTC-registered event contracts for non-sports outcomes; the Superior Court injunction was explicitly tied to sports-outcome contracts; the SJC has not ruled on extension to non-sports. | Commonwealth v. KalshiEX (MA Super. Ct., Jan. 2026) — no
docket, no division, no judge, no reporter; sourced via Wikipedia
[T4] |
| Qwen | Hostile and uncertain. A Massachusetts resident "cannot currently participate in these federally regulated markets without facing potential legal jeopardy." Suffolk County Superior Court injunction barring Kalshi sports contracts, requiring geofencing; 30+ active nationwide lawsuits with conflicting rulings. | Trade press and gambling-affiliate media [T4] |
Resolution: unsettled [T6]. Three
grounds, applied in the order the brief specifies.
First, on majority: two of three sources characterize the question as unresolved. Gemini stands alone in declaring it settled.
Second, on what the specific citation actually holds: Gemini's citation is the best-formed in the set, and it is retained — but KalshiEX LLC v. CFTC adjudicated whether the Commission could block election contracts under its own § 5c(c)(5)(C) review authority. It is not a holding that the Commodity Exchange Act preempts a state's application of its own gambling law to a state resident, and Gemini offers no authority that says so. The brief's rule is to prefer the more specific citation for the proposition it actually supports; a well-cited case does not extend to a proposition it did not decide. Gemini's "100% lawful" characterization is an inference stated as a holding.
Third, on chronology: the D.D.C./D.C. Cir. decisions are dated 2024. The Massachusetts Superior Court injunction, the Attorney General's suit, and the state-court actions catalogued across Arizona, Michigan, Minnesota, Nevada, New York, Ohio, Washington, and Wisconsin all post-date it, and MiniMax reports the CFTC filed an amicus brief in the Massachusetts Supreme Judicial Court — an action no one takes in a settled area. A 2024 federal decision does not settle a question that state courts were actively litigating in 2026.
The narrower question — non-sports contracts — is separately
unsettled. Every source that describes the Massachusetts
injunction describes it as reaching sports contracts
and requiring geofencing for sports markets (Sources: MiniMax,
Qwen). Whether it reaches economic, monetary-policy, or election
contracts is, in MiniMax's words, "unresolved by Massachusetts courts,"
and the SJC had not ruled at the research date. This section marks it
[T6]/unknown and declines to pick a side.
Practical posture for the reader. The conservative
reading — that Massachusetts treats event contracts as wagering under c.
23K until a court says otherwise — is the reading MiniMax itself adopts
as its conservative inference [T6]. Adopting it costs a
vehicle. Adopting Gemini's reading costs, potentially, a great deal
more, and rests on a preemption argument no cited authority makes. Given
that the sole stake at risk is USD 100 and the downside is a legal
exposure of unbounded size, the asymmetry strongly favors the
conservative reading — not because it is established, but
because it is the cheap error.
MiniMax's state-action lists are internally contradictory and
are not propagated. Its §A.3.2 enumerates nine states acting
against Kalshi (Arizona, Massachusetts, Michigan, Minnesota, Nevada, New
York, Ohio, Washington, Wisconsin); its §H.1 enumerates a
different nine states that the CFTC sued
(Arizona, Connecticut, Illinois, Kentucky, Minnesota, New Mexico, New
York, Rhode Island, Wisconsin), and lists that extraordinary posture as
settled with no docket, filing date, or court for any
of them. Six states appear in one list and not the other, and the two
lists describe opposite actor postures. Both lists are marked
[T6]. Qwen's "over 30 active lawsuits" is
trade-press sourced and time-sensitive [T4]. What survives
across all three sources is only the direction: prediction-market
legality was under active, multi-state, conflicting litigation at the
research date.
Authority. The Attorney General enforces
M.G.L. c. 93A, the Consumer Protection Act, which is
the office's principal hook here [T5] (Source: MiniMax,
Gemini — independently cited). MiniMax adds
M.G.L. c. 12 §§ 4L–5 for the office's general authority;
the range notation mixes a lettered and a numbered section and no
pincite is given for the specific power invoked, so that
citation is marked [T6] (Source:
MiniMax).
Enforcement record. Gemini cites a Robinhood
USD 7.5 million settlement (2024) for deceptive gamification as
the operative c. 93A precedent [T2] (Source:
Gemini). MiniMax reports Attorney General Andrea Joy
Campbell filed suit against KalshiEX in September
2025 alleging it accepted sports wagers without Massachusetts
authorization, leading to the January 2026 Superior Court injunction,
and describes the office as the primary state actor on prediction-market
enforcement [T4] — MiniMax's own text concedes the direct
MA AG URL was blocked at retrieval and the claim reaches it via press
reports and Wikipedia (Source: MiniMax).
Consequence for this experiment. Both sources that
address it agree that purely personal, self-directed algorithmic
execution against a broker's API is unimpeded by c. 93A
[T5] (Source: MiniMax, Gemini). Chapter 93A
reaches deceptive practices by a business toward
consumers; a principal trading their own USD 100 is on the protected
side of that statute, not the regulated side. The exposure flips only if
outputs are shared for compensation — see §11.9.
The Attorney General has taken no public position on
non-sports event contracts. MiniMax states this affirmatively
and marks it [T6], adding that the conservative inference
is that the office treats all event contracts as gaming under c. 23K
until a court rules otherwise (Source: MiniMax). That
characterization is adopted here, as [T6].
Qwen contributes nothing to this subsection — its tax coverage is entirely absent, which is a named CASINO requirement missed in full. Everything below comes from MiniMax and Gemini.
Short-term capital gains. Gains on capital assets
held one year or less are short-term and taxed at ordinary rates,
spanning 10%–37% [T5] (Source:
Gemini); the short-term definition sits at 26 U.S.C. §
1222(1) and the preferential long-term rate schedule at
26 U.S.C. § 1(h) [T5] (Source:
MiniMax). Within a 90-day experiment no position can reach
long-term treatment — a point worth stating because MiniMax's
own after-tax deliverable applies 15% long-term rates to its equity
baseline, which is impossible on a 90-day horizon and inflates the
apparent tax advantage of equities relative to every other vehicle.
Corrected to short-term, the equity path carries the same ordinary-rate
federal burden as the wagering characterization does.
Wash-sale rule. 26 U.S.C. § 1091(a)
disallows a loss on the sale of "stock or securities" where
substantially identical property is acquired within the 30-day window on
either side, with the disallowed loss added to the replacement basis
[T5] (Source: MiniMax, Gemini — independently
cited). Application by vehicle:
[T5].[T5].[T5] (Source: MiniMax,
Gemini — both state the exemption; MiniMax supplies the
subsection).[T5]
(Source: MiniMax). Single-sourced but flows directly from the
statutory text's own limiting language.[T5] (Source: MiniMax, Gemini).[T5] (Source: MiniMax).Section 1256 treatment. A § 1256 contract is marked
to market at December 31 and any gain or loss is split 60%
long-term / 40% short-term regardless of holding period, under
26 U.S.C. § 1256(a)(1) and §
1256(a)(3) [T5] (Source: MiniMax,
Gemini). Reporting is on Form 6781, carrying to
Schedule D [T5] (Source: Gemini). MiniMax
never names Form 6781 anywhere despite discussing § 1256
treatment across eight sections — a reader following MiniMax alone would
have no filing path. Gemini supplies it and it is adopted.
MiniMax attributes the 60/40 split to § 1(h)(6) in two
places while correctly attributing it to § 1256(a)(3)
elsewhere in the same report. § 1(h)(6) is dropped as
an internal contradiction resolved against itself. MiniMax's subsection
assignments within § 1256(g) — (g)(1)(A) used as if it were
the whole "regulated futures contract" definition,
(g)(7)(B) for the qualified-board/DCM prong,
(g)(1)(B) for foreign currency contract,
(g)(6) for nonequity index options adjacent to
(g)(3) for the same concept — are mutually
inconsistent within MiniMax's own tables and are
downgraded to [T6]. The concepts
(regulated futures contract; qualified board or exchange; nonequity
option) are correct and settled; the specific subsection letters are not
established by any source here and should be checked against the statute
before filing.
Section 1256 mark-to-market and the 90-day horizon.
MiniMax correctly notes that mark-to-market creates a tax liability
before cash is realized, then marks the consequence [T6]
and stops [T6] (Source: MiniMax). The unanalyzed
question matters: whether the 90-day window straddles December
31 determines whether mark-to-market bites at all. No source
states the experiment's start date. If the window closes before
year-end, the § 1256 position is closed and marked in the ordinary
course; if it straddles, an open position generates a December 31
recognition event on an unrealized gain, with no cash to pay it from.
[T6]/unresolved, and a scheduling decision the reader
controls.
Wagering losses. 26 U.S.C. § 165(d)
limits wagering-loss deductions to the extent of wagering gains
[T5] (Source: MiniMax). MiniMax reports an
amendment by Pub. L. 119-21, § 70114(a) further
limiting the deduction to 90% of losses
[T5] for the statutory citation (Source: MiniMax).
The effective date is unresolved. MiniMax states it two
incompatible ways within one report — "post-July 4, 2025" throughout,
and "the 2018–2025 period" once — and a single amendment cannot do both.
Marked [T6]. Since which tax year the 90%
limit bites determines the entire after-tax arithmetic of the wagering
branch, the effective date must be checked against the enacted text of
Pub. L. 119-21 § 70114 before any calculation is relied on. Wagering
gains are includible in gross income under 26 U.S.C. §
61 [T5] (Source: MiniMax).
The federal characterization of prediction-market proceeds is UNSETTLED. This is the most important tax finding in the section, and both sources that address it agree.
MiniMax states plainly: no IRS Notice, Revenue Ruling,
Private Letter Ruling, or regulation has resolved which characterization
applies to prediction-market event contracts, despite their
existence since 2021 [T5] — this is a statement about the
absence of authority, which is verifiable and which both
sources make independently. Gemini states: "Event contract status is
UNSETTLED — no IRS guidance" [T5] (Source: MiniMax,
Gemini). Gemini's own open-questions list names it first: no
binding IRS Revenue Ruling on binary prediction markets, § 1256
eligibility contested, formal tax counsel opinion needed.
The two competing characterizations:
[T6]. MiniMax's supports for this are the
CFTC's own "futures contract" framing of BTCPERP (release 9240-26, in
the unverified series) and a mark-to-market premise cited to
CEA § 5c(c)(4), which is a self-certification provision and
does not support it. With that cite dropped, Argument A's
mark-to-market premise is unsupported by any retained
authority. MiniMax concedes in its own text that CFTC
classification "is not binding for IRS purposes" — which is correct and
dispositive of the argument's weight.[T6].
MiniMax's only authorities are Rev. Rul. 54-339 — cited
with no bulletin reference, no holding quoted, no pincite — and the 1954
Code's § 4421 definition of "wagering transaction,"
which is a wagering excise-tax definition imported into
an income-tax loss provision with no stated bridge. Both are
marked [T6].Gemini adds a third possibility its JSON records but its markdown
does not develop: "Section 1256 Non-Equity Options vs Open
Contracts vs Ordinary Income," suggesting an open-transaction
treatment as a further branch [T6] (Source:
Gemini).
The decisive practical observation, and it belongs to
MiniMax: a retail trader cannot influence which form the
venue issues. The characterization is operationally in the venue's
hands [T5] (Source: MiniMax). MiniMax further
claims Kalshi's practice appears to treat contracts as § 1256 and issue
1099-B, and that Polymarket "similarly produces Form 1099-B" —
both claims are sourced to nothing, and MiniMax's own
unknowns list contradicts the second by recording Polymarket's form
practice as unknown. Both are marked [T6] and
should not be relied on. Gemini's parallel entry says "1099-B
or 1099-MISC," which is the honest answer.
Gambling-winnings reporting form — unresolved.
MiniMax asserts 1099-MISC Box 3 for the wagering branch
with no authority cited, and its own table concedes there is "no
specific slot." Form W-2G — the actual reporting mechanism for
gambling winnings, with its own issuance thresholds — appears zero times
in MiniMax's report and is absent from Gemini's. Gemini offers
"1099-B or 1099-MISC." No source establishes the correct form
for the wagering branch. Marked [T6]/unknown. This
is a real gap: MiniMax's own argument that the issued form determines
the tax outcome is built on an unsupported guess about which form that
is.
Crypto. Digital assets are property and capital
assets per IRS Notice 2014-21, with capital gain or
loss on disposition [T5] (Source: MiniMax).
Form 1099-DA applies for brokers in scope beginning
with 2025 transactions [T5] (Source: MiniMax).
MiniMax also cites Rev. Rul. 2023-14 for disposition
treatment while its own source list describes that ruling as addressing
staking rewards taxable at receipt — the ruling does
not support the proposition it is attached to, and it is
dropped (Source: MiniMax).
Foreign currency and CFDs. 26 U.S.C. §
988 treats foreign-currency gain or loss as ordinary by
default, with an election available [T5]; OTC retail FX and
CFDs are generally not 1099-B reportable [T5] (Source:
MiniMax). Out of scope at USD 100 but included for completeness of
the merged table.
Publications. IRS Pub. 550
(Investment Income and Expenses), Pub. 551 (Basis of
Assets), and Pub. 525 (Taxable and Nontaxable Income,
for the gambling branch) [T5] (Source: MiniMax;
accessed Aug 1, 2026). "IRS Pub. 5029," described by
MiniMax as an online digital-asset FAQ, does not correspond to a
recognizable publication and is dropped.
A 1099 may not be issued at all. MiniMax notes that
at USD 100 the account may fall below issuance thresholds, then hedges
that covered securities are typically reported regardless
[T6]. The "$20 for fractional shares" threshold it
cites corresponds to nothing else in its report and is dropped.
The operative rule is unchanged either way: income is reportable whether
or not a form arrives.
Massachusetts is where the divergence lives, and it is the one place in this section where the two tax-covering sources genuinely contradict each other on a question the reader will act on.
Conformity. Massachusetts personal income tax starts
from federal taxable income with adjustments under M.G.L. c. 62
§ 1, with gains taxed under M.G.L. c. 62 § 4
[T5] (Source: MiniMax, Gemini — both cite c. 62 §
4). MiniMax repeatedly characterizes conformity as "largely
automatic"; Massachusetts IRC conformity for personal income tax is
date-limited rather than rolling, which if correct would undercut that
premise and, with it, MiniMax's assumption that every federal
characterization flows through unchanged. The "automatic
conformity" premise is downgraded to [T6] and no
conclusion below rests on it alone.
Short-term capital gains rate — conflict resolved toward
Gemini. MiniMax gives Massachusetts capital gains as "5%,"
derived from an internally garbled passage ("Part A rate 5.05%," "Part B
9.5% on long-term gains," a "4% surtax" that does not reconcile with a
stated 9.5% tier) which MiniMax's own text flags [T6] —
"see [T6] for precise MA tax computation" — and then hard-codes as
settled into every after-tax figure it produces. Gemini gives
8.5%, reduced from 12.0%, citing Chapter 50 of
the Acts of 2023 and TIR 24-4
[T3] (Sources: Gemini vs. MiniMax; resolved toward
Gemini).
Resolution: 8.5%, at [T3]. Gemini names
a specific session law and a specific Technical Information Release;
MiniMax names a rate its own report declares unverified. The brief's
rule — prefer the more specific citation regardless of which source
stated it — points unambiguously to Gemini. The tier is
[T3] rather than [T5] because TIR 24-4 is an
agency technical release reaching this section through a single report
that did not retrieve mass.gov. Every after-tax figure in
MiniMax's report uses 5% and is therefore understated on the
Massachusetts component; do not propagate MiniMax's after-tax
arithmetic.
Gambling winnings rate. Gemini cites M.G.L.
c. 62 § 3(B)(a)(13) for gambling winnings taxed at
5.0% [T5] (Source: Gemini).
Single-sourced but specifically pincited to a subsection.
The Massachusetts gambling-loss trap — the divergence from federal treatment.
This is the finding the brief flags as diverging from federal law, and the two sources say opposite things.
[T6]. MiniMax's own headline heading for this
finding reads "settled but diverges from federal" while
its body says treatment is "identical to the federal
default" — the report contradicts itself in the space of one
finding. The sentence meant to deliver the divergence is, verbatim and
incomplete: "Massachusetts does not allow gambling losses as an
itemized deduction against ordinary income; the OBBBA change at the
federal level that permits 90% of losses for professional gamblers to
flow through to MA under conformity." The second clause has no
predicate. MiniMax never actually delivers a conclusion
here.[T5]/[T3] (Source:
Gemini). If prediction-market proceeds are recharacterized as
gambling, losses cannot offset wins at the state level
and the 5.0% Massachusetts tax applies to gross
winnings, not net.Resolution: Gemini's divergence finding is adopted, at
[T3], with an explicit verification flag.
The reasoning: a specific statutory subsection plus a numbered TIR
defeats a self-contradictory report whose own sentence on the point is
grammatically incomplete and which cites no Massachusetts authority for
its conformity claim. That is exactly the case the brief's "prefer the
more specific citation" rule is built for. Two cautions attach and both
are recorded. First, TIR 15-14 does not appear in Gemini's own
bibliography despite carrying a tier tag in its body — a
load-bearing citation that its own source list cannot corroborate.
Second, Gemini's JSON sources this same finding to a CPA firm's
marketing blog (camusocpa.com), which is
[T4] grey literature supporting the highest-stakes
unresolved claim in either report. The tier is therefore capped
at [T3] and the finding carries a standing instruction:
verify M.G.L. c. 62 § 3(B)(a)(18) and TIR 15-14 against primary text
before acting on it.
Why this matters more than the rate. Under the
gambling characterization the Massachusetts loss disallowance means a
trader can lose money on the year in aggregate and still owe
Massachusetts tax on every winning contract. That is not a rate
difference; it is a change in the tax base, and it interacts with the
unsettled federal characterization in §11.7 to produce genuine downside
asymmetry. Gemini quantifies the effect as raising the required
break-even win rate from 50.00% under capital-gains/§
1256 treatment (where losses offset dollar-for-dollar) to
51.66% if federal losses are itemized, or
57.80% if they are not [T6] (Source:
Gemini). The arithmetic checks internally, but the derivation is
not shown and the parameters (t_f = 0.22,
t_ma = 0.05, USD 1,000 cumulative turnover at 1:1 odds) are
assumptions rather than the reader's actual facts. Marked
[T6]. Note that Gemini's own executive summary
quotes only the worse figure (57.80%) without stating that it applies
solely to the non-itemized case; the itemized case, which is the more
common one, is 51.66%.
MiniMax's after-tax break-even deliverable is structurally
unsound and is not propagated. Its method applies a
1/(1−t) gross-up to gross proceeds when
tax falls on the gain, which produces a stated
"break-even" of USD 125 to USD 141 where the true post-tax break-even
exit is USD 100. Its Massachusetts computation additionally drops a term
in one scenario and double-counts state tax in another, so its own two
sections disagree on the same scenario. All MiniMax after-tax
figures are marked [T6] and excluded from this section's
tables (Source: MiniMax; error confirmed by MiniMax's own
digest).
Form divergence at the state level. Whichever
federal form the venue issues drives the Massachusetts result: 1099-B
Boxes 8–11 as a § 1256 contract yields Massachusetts capital-gain
treatment under conformity; 1099-MISC as wagering income yields
Massachusetts ordinary income under c. 62 § 1 [T6]
(Source: MiniMax). Combined with the loss-disallowance finding,
the state consequence of the wagering branch is materially worse than
the federal consequence alone suggests.
No MA DOR guidance on event contracts exists.
MiniMax records that no DOR Technical Information Release addressing
event contracts was located at retrieval [T6]; Gemini does
not claim otherwise. [T6]/unknown.
Massachusetts crypto. MiniMax cites M.G.L.
c. 169 for the state money-transmission regime, states
Massachusetts has not adopted a distinct crypto licensing regime and
does not prohibit crypto trading, and marks its own claim
[T6] for post-2025 developments [T6]
(Source: MiniMax). Single-sourced and self-flagged; retained at
[T6].
Both sources that address this agree on the answer, which makes it the most reliable conclusion in the section.
Personal use triggers nothing. A personal automated
analysis system with no external users, no shared outputs, no
compensation, operated entirely on the principal's own capital, is a
personal trading tool, not an investment adviser [T5]
(Source: MiniMax, Gemini — independently concluded). Gemini
states personal automated trading scripts for one's own account are
exempt under Investment Advisers Act § 202(a)(11)
[T5]. The statutory definition at 15 U.S.C. §
80b-2(a)(11) reaches a person who, for compensation,
engages in the business of advising others on securities
[T5] (Source: MiniMax) — a principal advising
themselves satisfies neither element.
The four elements that must all be present
(Source: MiniMax) [T5]:
The boundary moves the moment any of three things
happen (Source: MiniMax, Gemini)
[T5]:
At that point federal registration under 15 U.S.C. §
80b-3 is likely required, and Gemini adds that
distributing signals for compensation triggers Massachusetts
state RIA registration under M.G.L. c. 110A § 201
[T5] (Source: Gemini) — a specific state pincite
under the chapter this section adopted in §11.4, and the strongest
state-law citation in either report on this question.
The publisher's exclusion is narrow and its case authority
could not be verified. The Advisers Act excludes bona fide
publishers of regular financial publications at §
202(a)(11)(D) [T5], but MiniMax's own gloss is the
operative caution: the exclusion depends on compensation flowing from
subscription revenue rather than advisory fees, and
"sponsorships and affiliate kickbacks can pierce the exemption"
[T6] (Source: MiniMax).
MiniMax's sole case authority for this exclusion —
SEC v. Lowe, 7 F.4th 232 (2d Cir. 2021) — is dropped
entirely. MiniMax's own text concedes it was "not retrieved
this session," and the citation does not correspond to any decision the
report verified: court, reporter, volume, page, and year are all
unconfirmed. It is the only case citation in MiniMax's entire cluster
and it supports the report's answer to the registration question.
It is not propagated. The leading publisher's-exclusion
authority is a 1985 Supreme Court decision, but no source in this merge
supplied a verified citation to it, so the case-law scope of the
publisher's exclusion is marked [T6]/unknown here.
The statutory exclusion at § 202(a)(11)(D) stands at [T5];
its judicial construction does not.
The federal/state registration threshold is
unresolved. MiniMax cites § 203A(a)(1)(B) for a
USD 100 million AUM line between state and federal registration, then
disclaims it in the same sentence — "the Dodd-Frank Act changed this in
some cases; check current SEC rules." A threshold stated with a
subsection cite and immediately withdrawn is not a finding.
Marked [T6]. At any plausible scale for
this experiment the state channel is the operative one regardless.
Massachusetts registration mechanics. MiniMax names
"Form MA IA" twice as the Massachusetts
investment-adviser registration form, with no URL and no
issuing-authority confirmation, while pairing it with CRD — which is the
correct system. "Form MA IA" is not a recognized filing and is
dropped; state IA and IAR registration runs through the
IARD/CRD systems, but no source in this merge cites the specific form,
so the filing vehicle is marked [T6].
MiniMax's de minimis threshold — fewer than five Massachusetts clients
in twelve months with no holding out — carries the c. 110H chapter
defect corrected in §11.4 and is marked
[T6] pending verification against c. 110A.
Other registration triggers, none of which apply
here (Source: MiniMax) [T5]: a direct
public offering of partnership interests in a personal trading vehicle
would trigger Securities Act registration absent a Rule 506 safe harbor;
operating a hedge-fund-style vehicle for others would trigger Investment
Company Act registration absent the § 3(c)(1) or § 3(c)(7) exclusions,
neither of which fits this fact pattern.
The one-sentence answer. A self-directed,
personal-use automated analysis system with no published outputs and no
compensation from any source triggers no SEC and no
Massachusetts investment-adviser registration
[T5]. Publish it, or take a dollar for it, and that
changes.
Certainty column uses: settled (two sources or one
specific statutory citation with no contradiction),
contested (sources disagree, resolved),
unsettled (no governing authority exists),
unknown ([T6], cannot be determined from
sources).
| Vehicle | Federal characterization | Federal forms | Wash-sale applies | MA characterization | MA loss deductibility | Certainty | Source |
|---|---|---|---|---|---|---|---|
| Equity / ETF, held ≤ 90 days | Short-term capital gain; ordinary rates 10%–37% (§ 1222(1); § 1(h)) | 1099-B, Form 8949, Schedule D | Yes (§ 1091(a); reported Box 1g) | Short-term capital gain, 8.5% (c. 62 § 4; Ch. 50 Acts of 2023; TIR 24-4) | Full dollar-for-dollar offset against capital gains | settled federally; contested→resolved on MA rate | MiniMax, Gemini |
| Listed equity option (long call/put) | Capital asset; short-term at this horizon; treatment on exercise per § 1234 | 1099-B, Form 8949, Schedule D | Yes where substantially identical | Follows federal; MA capital gain at 8.5% | Capital loss against capital gains | settled | MiniMax, Gemini |
| Broad-based index option / nonequity option | § 1256 contract; 60% LT / 40% ST regardless of holding period (§ 1256(a)(3)) | Form 6781, Schedule D; 1099-B Boxes 8–11 | No (§ 1256(f)(5)) | MA capital gain; 8.5% on the short-term portion | Full mark-to-market offset | settled (subsection letters within § 1256(g)
[T6]) |
Gemini, MiniMax |
| Regulated futures contract (CME futures; KalshiEX BTCPERP) | § 1256 contract; 60/40; mark-to-market at Dec 31 | Form 6781, Schedule D; 1099-B Boxes 8–11 | No (§ 1256(f)(5)) | MA capital gain under conformity | Capital loss against capital gains | settled for CME futures;
[T6] for BTCPERP (rests on unverified CFTC
release 9240-26) |
MiniMax, Gemini |
| Event contract (Kalshi / Polymarket / ForecastEx) — Argument A, § 1256 | Regulated futures contract on a DCM → 60/40 | Form 6781; 1099-B Boxes 8–11 | No (§ 1256(f)(5)) | MA capital gain under conformity, 8.5% short-term portion | Capital loss against capital gains | UNSETTLED — no IRS Notice, Rev. Rul., PLR, or reg addresses this vehicle class | MiniMax, Gemini |
| Event contract — Argument B, wagering | Ordinary income under § 61; losses limited by § 165(d) (90% limit
per Pub. L. 119-21 § 70114(a); effective date
[T6]) |
Form unknown [T6] — MiniMax asserts
1099-MISC Box 3 without authority; Gemini says "1099-B or 1099-MISC";
W-2G is named by neither |
No (not "stock or securities") | MA ordinary income; gambling winnings 5.0% (c. 62 § 3(B)(a)(13)) | Losses NOT deductible — MA disallows gambling-loss deductions for non-MA-licensed wagering (c. 62 § 3(B)(a)(18); TIR 15-14). Tax applies to GROSS winnings. This is the federal/MA divergence. | UNSETTLED federally; MA divergence
contested→resolved toward Gemini at [T3], verify
before acting |
Gemini (MA divergence); MiniMax (federal branch) |
| Event contract — Argument C, open transaction | Named as a third possibility, not developed by any source | — | — | — | — | [T6] / unknown |
Gemini (JSON only) |
| Spot crypto | Property / capital asset (Notice 2014-21); short-term at this horizon | Form 1099-DA (brokers in scope, 2025 transactions forward); Form 8949, Schedule D | No — not "stock or securities" under § 1091(a) | Follows federal; MA capital gain at 8.5% short-term | Capital loss against capital gains under conformity | settled federally; MA rate at
[T3] |
MiniMax |
| OTC retail FX / CFD | § 988 ordinary by default; election available under § 988(a)(1)(B) | Generally not 1099-B reportable | No (ordinary under § 988) | MA ordinary income under conformity | Limited | settled; out of scope at USD 100 | MiniMax |
| Short sale of stock | Short-term capital gain or loss (§ 1233) | 1099-B, Form 8949, Schedule D | Yes (§ 1091(a), § 1091(e)) | Follows federal | Capital loss against capital gains | settled; infeasible at USD 100 (requires margin) | MiniMax |
Three notes on this table.
First, the § 1256 row for BTCPERP and both event-contract rows are the only places in this experiment where the tax outcome is genuinely indeterminate, and the indeterminacy is not the reader's to resolve — the venue's choice of reporting form drives it.
Second, the "Certainty: settled" values in the equity, option, and crypto rows describe the federal characterization only. The Massachusetts 8.5% rate underlying all of them rests on a single source's citation to TIR 24-4, unretrieved.
Third, MiniMax produced three mutually inconsistent tax tables with different row sets, different certainty values for the same question, and different source columns — one of them labels the § 1256 branch "settled" while another labels the identical question "unsettled." The table above resolves that internally: the § 1256 mechanism is settled; its application to event contracts is not.
Stripping the citations out, five things bind:
[T5].[T4]/[T5].[T5].[T6].[T5].Each entry: conflicting claims → supporting reports → resolution and reason. Entries marked [DOWNGRADED] were moved to unsettled/unknown out of caution rather than by picking a side.
[T6]/unsettled. MiniMax's claim reaches
the merge only through Wikipedia citing a notice never retrieved, and
MiniMax's own text concedes brokers may still enforce the floor.
Cautious default applied. Operationally moot — margin is unavailable at
USD 100 either way.[T6],
unresolved. MiniMax's is wrong on its face (BSA CIP rules are
in 31 CFR, not 17 CFR); Gemini's Part 1020 is the banks subpart, not
broker-dealers. Neither source cites the broker-dealer subpart. The
requirement is settled; the section number is not established.
Not resolved by picking the less-wrong option.[T5], pincite
downgraded to [T6]. The carve-out flows from the
absence of a recommendation, not from an express exclusion.[T4]; FINRA 2360 governs account
approval and diligence [T5]. Presenting the ladder as a
regulatory requirement would be wrong.[T6]. Four incompatible
figures, none with an underlying, strike, expiry, or quote date.[T6]/unsettled. Three grounds: (1) 2 of 3
say unresolved; (2) Gemini's citation — the best-formed in the merge and
retained as such — holds that the CFTC could not block those
contracts, not that the CEA preempts state gambling law
against a state resident, and Gemini offers no authority for the
extension; (3) the 2024 federal decisions predate the 2026 state
actions, and the CFTC's amicus filing in the MA SJC is itself evidence
the question was live. This is the single most consequential downgrade
in the section.[T6]/unknown. MiniMax's own text says
"unresolved by Massachusetts courts."[T6], including
9240-26, which anchors the BTCPERP § 1256 conclusion.[T6].[T6]/unverified. A
venue whose registration is unverified should not carry tax-treatment
rows.[T4].[T6].[T6].[T4], not load-bearing; two separately
numbered statutory schemes conflated.[T6]. The c. 93A hook is
[T5] and is corroborated by both sources.[T6] and
which contains three irreconcilable rate figures. Gemini: 8.5%, reduced
from 12.0%, citing Chapter 50 of the Acts of 2023 and TIR 24-4. →
8.5% at [T3]. Specific session law and
numbered TIR beat a self-flagged-unverified approximation. Consequence:
all of MiniMax's after-tax figures understate the MA
component.[T3] with a standing verification
flag. Specific subsection plus numbered TIR beats an incomplete
sentence. Tier capped at [T3] because TIR 15-14 is
absent from Gemini's own bibliography and Gemini's JSON sources the same
claim to a CPA marketing blog [T4]. Verify against
primary text before acting.[T6]. No conclusion in this section rests
on it alone.[T6]. A single amendment cannot do both.
The statutory citation (Pub. L. 119-21 § 70114(a)) is retained
[T5]; the effective date must be verified before any
after-tax calculation.[T5]; specific subsection letters
[T6].[T6]/unknown. Consequential: MiniMax's own
thesis is that the issued form determines the outcome, so the thesis
rests on an unsupported guess.[T6], self-contradicted.[T6]. These are the only authorities
offered for Argument B.SEC v. Lowe, 7 F.4th 232 (2d Cir. 2021).
[DOWNGRADED] MiniMax's only case authority, cited for the
publisher's-exclusion scope on which its registration answer turns, and
conceded by MiniMax as "not retrieved this session." → Dropped
entirely. The statutory exclusion at § 202(a)(11)(D) stands
[T5]; its judicial construction is
[T6]/unknown because no source supplied a verified
case citation.[T6].[T6] since no source cites the actual form.1/(1−t) gross-up to
gross proceeds when tax falls on the gain, producing "break-evens" of
USD 125–141 where the true post-tax break-even exit is USD 100; one
scenario drops a term and another double-counts MA tax, so its own
sections disagree on the same scenario. → All figures
[T6], excluded from this section's tables.
Gemini's break-even win-rate figures (50.00% / 51.66% / 57.80%) are
arithmetically sound but rest on undisclosed parameter assumptions and
are also [T6].Content gaps by source. Qwen omitted every tax authority — no IRS, no MA DOR, no capital-gains treatment, no wash-sale rule, no Section 1256, no forms — and never names the MA Securities Division, the MA Gaming Commission, or the MA Attorney General as agencies, referring only to "Massachusetts regulators." Both other sources cover all eight authorities. Gemini's JSON appendix omits the IRS, MA AG, and MA DOR (5 of 8 authorities present), though its markdown covers all eight; a consumer reading only Gemini's machine-readable half would lose the entire tax analysis. MiniMax covers all eight but handles the IRS and MA DOR as one-line stubs under its authority map, and omits Form 6781, Form W-2G, and Form 1099-K entirely — the first two being the reporting forms for the two characterizations its whole analysis is organized around. Two authorities are named by no source and are flagged as gaps: the NFA (relevant to any FCM intermediating DCM access) and SIPC / CFTC Part 190 customer-protection regimes (relevant to what happens to the balance if a venue fails — a live question given the litigation posture described above).
This section merges the three independently-commissioned research reports (Qwen, Gemini, MiniMax) into one ranking and one numeric conclusion. Where the three disagree, the resolution rule is stated and the losing figure is recorded rather than quietly dropped. Every substantive claim carries an evidence tier. The section refuses to be encouraging: where the merged evidence supports a negative finding, the negative finding is the headline.
The three source reports used the same column names with three different meanings, which is the primary reason their numbers appeared to diverge more than they actually do. The merged table fixes one definition per column. (Source: MiniMax §1.1; Gemini Table G; Qwen Table G)
[T1] for the framework,
[T6] for every number in the column.[T6]null in every row, with cause. Neither
MiniMax nor Gemini derives a first-passage time distribution
anywhere; both derive only terminal-return distributions and then state
days-to-target and interquartile ranges to two significant figures.
MiniMax's 45–75 / IQR 20–55 and Gemini's 65 / [42, 88] have no visible
construction. Filling these cells would be fabrication;
null with a stated reason is the honest entry. Gemini's
Martingale row is the reductio: P(reach) = 0.00 alongside
"median days 12, IQR [4, 22]" — a strategy that reaches the target with
probability zero has no distribution of days-to-target.
[T6][T5] for published schedules, [T6] for the
trade-count assumption.[T6][T1] peer-reviewed and
replicated; [T2] peer-reviewed, unreplicated or contested;
[T3] working paper/preprint; [T4] grey and
industry literature; [T5] primary regulatory, statutory,
exchange or vendor documentation; [T6] author inference,
unverified.Inclusion rule for the ranking, stated once and
applied consistently: a strategy is ranked if it is (a) executable at
USD 100 by a Massachusetts-resident retail participant and (b) not
documented in the literature as defeated. Everything else goes to the
excluded block with its figures preserved. All three sources violated
their own inclusion rules — MiniMax ranked rows falling below its stated
1% cutoff and ranked a cash-secured short put that requires roughly USD
2,000 of collateral; Qwen ranked prediction-market speculation second
while concluding twice that it is legally inaccessible to the subject;
Gemini ranked Martingale sizing, which it had just proven produces
certain ruin. [T6]
Ranking key: merged central estimate of P(reach), descending; ties broken by P(ruin), ascending.
All figures are merged across the three reports per the resolutions in §12.5. Ranges communicate genuine forecast uncertainty, not hedging.
| Rank | Strategy | P(reach $200) | P(ruin ≤ $25) | Median days to target | IQR of days | Total friction drag | After-tax EV (Δ over 90d) | Evidence tier |
|---|---|---|---|---|---|---|---|---|
| 1 | Spot crypto held outright (BTC or comparable high-vol major), weekly rebalance, no leverage | 0.03–0.08 (central 0.05) | 0.40–0.60 | null — no FPT distribution derived by any source |
null |
0.8–3.0% | −$12 to −$29 | [T1] realized-return distribution; [T5]
venue fee schedules; [T6] probabilities |
| 2 | High-beta single-name equity, full $100 position, 60–90 day hold, unlevered | 0.02–0.06 (central 0.04) | 0.35–0.55 | null |
null |
0.02–2.0% | −$10 to −$20 | [T1] cross-sectional vol; [T6]
name-specific probabilities |
| 3 | Informed CFTC event-contract trading (Kalshi / ForecastEx), calibration-based favorite selection, maker limit orders — MA access contested | 0.02–0.05 (central 0.035) | 0.30–0.45 | null |
null |
3.6–15% | −$8 to −$18 | [T1] favorite-longshot bias, market calibration;
[T5] fee schedules, Commonwealth v. KalshiEX;
[T2]/[T6] retail informed-trader
profitability |
| 4 | Long premium directional options, single-leg (incl. full-stake bold play and 0DTE variants) | 0.01–0.08 (central 0.02, disputed) | 0.55–0.92 | null |
null |
2.5–12% | −$20 to −$58 | [T3] 0DTE retail literature; [T4] industry
loss data; [T6] probabilities |
| 5 | Post-earnings announcement drift (PEAD) on commission-free fractionals, top-decile SUE | 0.005–0.02 (central 0.01) | 0.15–0.30 | null |
null |
1–12% | −$12 to −$25 | [T1] drift effect (Bernard & Thomas 1989);
[T1] post-publication decay (McLean & Pontiff 2016);
[T6] retail feasibility |
| 6 | Cross-sectional / ETF momentum, long-only at $100, weekly–monthly rebalance | 0.005–0.02 (central 0.01) | 0.20–0.35 | null |
null |
0.25–10% | −$8 to −$21 | [T1] effect and 58% post-publication decay;
[T6] net-of-friction feasibility |
| 7 | Short-term reversal / mean reversion, liquid equities, 1–5 day holds | 0.005–0.02 (central 0.01) | 0.25–0.40 | null |
null |
5–12% | −$8 to −$18 | [T1] Jegadeesh (1990) effect; [T6] net
feasibility |
| 8 | CFTC event-contract longshot lottery play (YES at $0.05–$0.10, multi-contract) | < 0.01 (central 0.005) | 0.70–0.90 | null |
null |
5–25% | −$20 to −$35 | [T1] favorite-longshot bias operates against the buyer;
[T5] fee formula; [T6] probabilities |
Row 3 carries a jurisdictional condition. Two of
three sources hold that Massachusetts access to CFTC event contracts is
contested or foreclosed: MiniMax cites the January 2026 Commonwealth
v. KalshiEX LLC Superior Court injunction with MA SJC
federal-preemption review pending and states plainly that if the
question does not resolve favorably, its own row 1 (this merged row 3)
collapses; Qwen concludes an MA resident "cannot currently participate
in these federally regulated markets without facing potential legal
jeopardy." Gemini's contrary claim — that Kalshi and ForecastEx are
"100% lawful venues for MA residents" — rests on federal preemption
under CEA § 2(a)(1)(A) without engaging the state injunction at all.
[T5] The merged position: row 3 is conditional on the MA
question resolving favorably or on the participant not residing in
Massachusetts. The headline in §12.4 survives row 3 collapsing
entirely, because the ceiling of the merged band is driven by
spot crypto (row 1), not by event contracts. [T6]
Row 4 carries an unresolved band. MiniMax puts its
0DTE structure below 0.01, Qwen puts deep-OTM option buying below 0.05,
Gemini puts long premium options at 0.08. None models it. The mechanical
argument runs the other way from the rest of the table — under
Dubins–Savage, bold play maximizes P(reach) in a subfair game, and a
single out-of-the-money call doubles on a far smaller underlying move
than the underlying's own doubling requires [T1]. That is
precisely why this row also carries the table's worst P(ruin) and its
worst after-tax expected value. The band is wide because the uncertainty
is real. [T6]
These are not ranked because they fail the inclusion rule. Merged non-redundantly across all three reports. (Source: MiniMax §1.4, Gemini §"Strategies That Do Not Work", Qwen §"Strategies Demonstrably Ineffective")
| Strategy | Reason excluded | Tier |
|---|---|---|
| Volatility-risk-premium harvest via cash-secured short put | Requires ~$2,000 collateral; short-option margin blocks a $100 cash account. Two of three sources grade it capital-infeasible despite VRP itself being real (Carr & Wu 2009) | [T1] effect; [T5] margin rules |
| Buy-and-hold broad index / SPY | P(2× in 90 days) < 0.1% at ~15–25% annualized vol; ln 2 sits 2σ–4σ above the mean | [T1] |
| T-bills / money market | Short rates imply < 2% over 90 days; horizon-infeasible by arithmetic | [T1] |
| 2×/3× leveraged and inverse ETFs held > 1 day | Volatility compounding drag exp(½(L−L²)σ²t); a 3× ETF
on a flat index at σ = 25% loses ≈ 4.6% over 90 days from path
alone |
[T1] mechanism |
| Retail day trading, any frequency | Taiwan population study: 99% net unprofitable, <1% show repeatable skill. Brazil, 19,642 traders over 300 days: 97% lost money, 1.1% earned above minimum wage, 0.1% above $300/day | [T1] |
| Penny / OTC / pink-sheet stocks | Bid-ask spreads 10–50% of share price; toxic convertible dilution; pervasive pump-and-dump | [T1] |
| Social-media, meme and sentiment-only signals | Sentiment metrics lag price; retail entry clusters at peak sentiment as institutional reversion begins | [T1] |
| Naive ML on price series without purged cross-validation | Overlapping forward-return labels leak; in-sample Sharpe > 4.0 collapses to zero or negative live | [T1] |
| Copy-trading and paid signal services | Adverse selection, execution latency, provider-first execution; seller's incentive is subscription revenue | [T2]/[T4] |
| Technical-analysis pattern rules | 15,000+ rules on ~100 years of Dow data under White's Reality Check and FDR correction: zero survive out-of-sample after 5–10 bps costs | [T1] |
| Martingale / anti-martingale / progressive sizing | From $100 with $1 base bet, bankruptcy arrives on the 7th consecutive loss; P(ruin) → 1 as trial count grows | [T1] |
| Index-reconstitution arbitrage | Effect ~1.5–3% one-time, Russell rebalance is annual (June); requires close-of-reconstitution institutional execution | [T1] effect; [T6] feasibility |
| Merger arbitrage | Needs $500–$1,000 minimum, rapid settlement, ability to short the acquirer | [T1] effect; [T6] feasibility |
| Regulated crypto nano futures | Minimum position $20–$50 (20–50% of the stake); liquidation risk on a $100 balance | [T5] |
| Polymarket | US IP geo-blocking under the 2022 CFTC consent order; federal authorization settled November 2025 via the QCEX acquisition but MA access remains contested | [T5] |
Across the full surveyed universe, the realistic probability
that USD 100 becomes USD 200 within 90 days under the best-supported
approach is 1% to 8%, with a central estimate of 3%. The corresponding
probability of an experiment-killing loss — terminal wealth of USD 25 or
less — is 45% to 75%, with a central estimate of 60%. No approach in the
surveyed universe carries positive expected value after costs and taxes
at USD 100 scale. [T6] for the probabilities;
[T1]/[T5] for the components they rest on.
Why the aggregate central (0.03) sits below Table G's rank-1 central (0.05). The two are different quantities and the gap is deliberate. Rank 1's central is the midpoint of a single row's band; the aggregate central is the figure defensible across the surveyed universe after two downward adjustments that apply to the ranking as a whole. First, rank 1's band is the least coherently derived in the table — its own supporting section states the number three incompatible ways (Conflict 5), so its midpoint carries less weight than its position suggests. Second, the retail reference-class floor of P(reach) < 1% pulls the universe-wide estimate down toward the bottom of the published band, and nothing in the merged evidence justifies letting a single crypto row set the aggregate. Both adjustments run in the same direction, and it is the honest direction: 0.03 is the number to quote, and 0.05 is the most optimistic single row rather than a summary of the universe.
Note on the ruin figure. The 45–75% band and its 60%
central are universe-wide aggregates, not
rank-1-specific — rank 1's own P(ruin) band is 0.40–0.60. They also
measure an experiment-killing loss (≤ USD 25), not literal total loss.
Genuine total loss is strategy-dependent and asymmetric: for the top
three ranked rows, which are unlevered spot positions, P(terminal wealth
= USD 0) is close to zero, and the modal bad outcome is a partial loss
in the USD 40–85 range rather than a wipeout. For rows 4 and 8 — long
premium options and event-contract longshots — total loss is real and
probable, at 0.55–0.92 and 0.70–0.90 respectively. The approaches that
most reliably destroy the entire stake are the ones that most resemble a
lottery ticket. [T6]
That third sentence is the headline finding, and it is the only one
of the three on which all three independent reports agree without
qualification. Qwen: "The evidence does not support the existence of any
approach that carries a positive expected value after accounting for the
inevitable frictions and the statistical realities of market
efficiency." Gemini: "NO strategy carries a positive risk-adjusted
expected value (E[W₉₀] < $100)." MiniMax: "every strategy in Table G
has negative expected value after taxes and friction at USD 100 scale."
Three reports built from three different literatures, three different
vehicle universes and three different modelling approaches converged on
it. It is the most robust conclusion in this report.
[T1]
Three mechanisms produce it, and they compound rather than substitute. (Source: MiniMax §2.2, corroborated by Gemini §11 and Qwen)
[T5]/[T6][T5] statutes;
[T6] the arithmetic[T1]On the expected-value convergence, one merge finding is worth
surfacing. The three reports diverge roughly fourfold on
P(reach) but converge tightly on after-tax expected value once Gemini's
terminal-wealth figures are converted to deltas: Gemini's top row
terminal USD 88.50 is a −USD 11.50 delta against MiniMax's −USD 8 to
−USD 18 for the same strategy; Gemini's PEAD row terminal USD 82.10 is
−USD 17.90 against MiniMax's −USD 12 to −USD 25. The agreement on the
sign and magnitude of expected loss, reached independently, is stronger
evidence than either report's probability estimate.
[T6]
Asymmetric uncertainty. The downside is open: an
adverse 2026 liquidity event raises the P(ruin) ceiling without a
symmetric effect on P(reach). The upside is bounded: no plausible
alternative assumption in the merged evidence raises P(reach) above
roughly 15%, and the one source that exceeded that bound did so without
any derivation (see Resolved Conflicts, Conflict 1).
[T6]
One caveat this section will not bury. The portion
of the merged band above ~1% is its least-supported part. Cluster 5's
empirical floor is P(reach) < 1%; the band's central estimate of 3%
sits above that floor on the strength of documented edges in the ranked
strategies — while §12.2's own EV column concludes those same edges are
consumed by friction and tax. The defensible reconciliation is that the
day-trading reference class is not the same reference class as the
best-supported approaches, so a modest uplift is warranted. But that
uplift is an inference, not a measurement, and readers should treat 1–2%
as the better-anchored end of the band and 8% as the end that depends
most heavily on a single incoherently-derived crypto row. Stated
directly: the true value is more likely to sit near the bottom
of the published band than the top. [T6]
What would move the headline — none of it is in
evidence as of the research date: a documented positive edge with a
>90-day track record at comparable scale; IRS resolution of event
contracts to § 1256; a regime of persistent, exploitable
prediction-market mispricing at accessible price points; or a strategy
with net pre-tax expected return exceeding roughly 8% per week for
twelve consecutive weeks with out-of-sample validation. The published
literature identifies no such strategy. [T6]
What would not move it: the negative-EV finding
survives every sensitivity the three reports tested, including full
collapse of the top-ranked event-contract row on jurisdictional grounds
and either direction of the IRS characterization question.
[T6]
The premise, on which all three sources agree in substance: under the
modal outcome — failure to reach USD 200 — the experiment produces
something generalizable if and only if it is run as an
instrumented study rather than as a trade log. Without
pre-registration, every observed outcome can be rationalized after the
fact, and the result is a narrative rather than a record. That is what
most retail trading logs are, and why most of them have no epistemic
value. [T1] methodology; [T6] application
Pre-register before the first trade. A dated,
immutable plan specifying strategies, sizing rules, stopping rules and
success criteria. This is the single highest-value artifact for
separating evidence from narrative. [T1]
Log per decision (19 fields):
decision_id; decision_timestamp_utc;
instrument; thesis (≤280 chars, ex ante);
probability_estimate_pre;
probability_estimate_post_resolved;
outcome_realized {full, partial, scratch, ruin-step};
position_size_usd;
position_size_rule_predicted;
position_size_rule_deviation_pct;
decision_latency_ms;
friction_drag_realized_usd;
friction_drag_assumed_usd; tax_realized_usd;
tax_assumed_usd; data_sources_consulted;
was_purged_validation_used;
confidence_in_thesis_pre (1–5);
notes_post_mortem. (Source: MiniMax §3.1)
Compute daily and weekly: per-trade Brier score
(p − outcome)²; Murphy's score
decomposition BS = REL − RES + UNC, which
separates calibration from discrimination and is strictly more
informative than a raw Brier [T1] (Source:
Gemini); Brier skill score against climatology; calibration slope
and intercept from outcome ~ logit(p); calibration drift
over a rolling 30-trade window; position-sizing rule adherence;
decision-latency distribution against outcomes; realized-versus-assumed
friction drag; after-tax versus pre-tax P&L ratio; per-strategy
attribution; cumulative realized EV.
Audit execution quality separately from strategy
quality — the basis-point gap between backtest fills at
mid-price P_mid and live fills P_fill. This is
the input to a realistic micro-capital friction model and is rare even
in published retail trading studies, which makes it the component most
likely to be independently useful. [T6] (Source:
Gemini, MiniMax)
Validate the pipeline as a deliverable in its own
right — WebSocket latency, connection dropouts, query execution
times, schema-validation errors. If the capital objective fails,
verified infrastructure is a real asset. [T6] (Source:
Gemini)
Pre-register stopping rules with alpha-spending
logic borrowed from clinical trials (Lan–DeMets,
O'Brien–Fleming): abandon on a cumulative loss floor (e.g. USD 50);
abandon if the Brier skill score against climatology is worse than
climatology at the 30-trade mark under sequential-testing correction;
abandon if the calibration slope is statistically distinguishable from 1
at the 30-trade mark under Benjamini–Hochberg correction across
pre-registered criteria; terminate any sub-strategy at a pre-set
drawdown. Record whether each rule fired and whether it was
obeyed — the second is the more informative datum.
[T1]
Know the power ceiling before starting. Two
independent results, from two sources, agree: 50 trades gives a realized
Sharpe of 1.0 a 95% confidence interval of roughly 0.4 to 1.6 (Lo 2002)
[T1]; and reaching t ≥ 3.0 over 90 trading
days requires SR_daily ≥ 3/√90 = 0.3162, i.e. an
annualized Sharpe of ≥ 5.02 — a figure that essentially
does not exist in unlevered retail asset classes [T1]
(Source: Gemini, MiniMax; arithmetic independently verified in both
digests). Therefore: a negative realized Sharpe combined with a
calibration slope significantly below 1 rejects both
the strategy and the belief system. A positive realized Sharpe
confirms nothing — it is consistent with luck under
almost any plausible edge magnitude. The experiment cannot in principle
validate a strategy at conventional confidence levels.
The salvage plan does not convert a low-probability experiment into a high-probability one. It converts a trade log into an instrumented study, which is the only conversion available. The strongest available learning is methodological, not strategic. A reader who finds that unsatisfying is correct to find it unsatisfying.
[T6] author-inference count = 0 while the entire
quantitative spine is author-modeled; (c) its top-ranked strategy
depends on an uncited assertion that Kalshi maker orders are free. It
also breaches MiniMax's own falsifiable ceiling ("no plausible
alternative assumption raises P(reach) above ~15%"). Qwen and MiniMax
agree within the 1–8% band; Gemini is the unsupported outlier and is not
averaged in.1 − P(reach) (0.18/0.82, 0.12/0.88, 0.09/0.91, 0.05/0.95,
0.08/0.92, 0.01/0.99, 0.00/1.00), which would mean every non-doubling
outcome is total loss — absurd for PEAD on commission-free fractionals.
Resolved: column discarded wholesale; merged P(ruin)
figures derive from MiniMax as corrected in Conflict 3.null in every row. Reason: neither source derives
a first-passage time distribution anywhere — both derive
terminal-return distributions only. These are the most
fabrication-shaped numbers in either report. Gemini's Martingale row (P
= 0.00 with median 12 days, IQR [4, 22]) demonstrates the column is not
tracking anything real.Merged headline figures (for direct lift into the executive summary and JSON appendix):
best_p_reach_200_in_90_days: 0.03
(band 0.01–0.08) — universe-wide central estimate under the
best-supported approach. Table G's single best row (spot crypto)
centrals at 0.05; see §12.4 for why the aggregate sits below it.corresponding_p_ruin: 0.60 (band
0.45–0.75) — universe-wide aggregate, and
"ruin" means terminal wealth ≤ USD 25 (experiment-killing loss),
not literal total loss. Literal total loss (terminal wealth =
USD 0) is near zero for the top three ranked rows, which are unlevered
spot positions, and is 0.55–0.92 for long premium options and 0.70–0.90
for event-contract longshots. Any downstream text labelling this figure
"probability of total loss" must use the ≤ USD 25 definition or
substitute the row-specific figure.positive_expected_value_exists: false
— the only unanimous finding across all three source reports.Genuinely unsettled in primary law, not merely under-researched:
Would require paid data or primary research to resolve, not achievable from AI-generated deep-research synthesis:
median_days_to_target field in this report's JSON appendix
is null for this reason, not because the value is unknown
but because no source computed it. This is flagged in this report's own
adversarial self-critique as its single weakest point.unsettled/[T6] precisely because no source
supplied one.Structural limitation of the merge process itself, documented in full in the adversarial self-critique below: majority-rule reconciliation across three AI-generated research reports assumes disagreement is informative and agreement is corroborating. Both assumptions can fail simultaneously if the three underlying tools share correlated failure modes (similar training data, similar retrieval patterns, similar fabrication triggers) — in which case agreement reflects shared error rather than independent verification, and this merge's conflict-resolution rules would not catch it. One instance of correlated risk was caught this pass, by chance, because the fabricated citation was single-sourced rather than agreed upon; a reader should not assume every fabrication in the underlying reports was necessarily this legible.
Every distinct citation appearing across Sections 3–12 of this merged report, deduplicated and organized by evidence tier. Nothing appears here that does not appear in those sections.
Four conventions govern this list and are applied without exception.
[T1] on content, [T6] on the
identifier). Each citation appears once, filed under
the highest tier assigned to it anywhere in Sections
3–12, with the split recorded in the trailing note. No citation
is duplicated across tier blocks.Scope. Python package documentation (Section 9, Table E) and data-source documentation (Section 10, Table F) are included, both having been presented as sources with consulted URLs. Packages appearing only on Section 9's unmaintained/avoid list are excluded — they are subjects under discussion, not sources cited. Within Tier T5, entries are grouped into legal/regulatory authorities and documentation sources, each alphabetical, because a single alphabetical run mixing statutes with package names is unusable.
[T1] on the mathematics in Section 3, [T6] on
the sourcing there (never reaches its citing report's own source list);
[T2] as a monograph in Section 6. (Sections 3,
6)[T1] for the effect, [T2] for its
independence. Corrects a source citation naming Israel as third author
and AQR as venue. (Section 5)[T1] on content;
[T6] on the DOI string, unverified in this
merge. A speculative alternative DOI (10.1111/0022-1082.00341)
printed in one source was excluded as a fabrication risk and must not be
reintroduced. Graded [T2] where Section 7 uses it for the
chart-watching channel. (Sections 3, 7, 12)[T6] on the DOI
string, uncited in its reporting source's own body.
(Sections 3, 7)[T6], attribution explicitly in
doubt: the work is principally a theoretical and simulation
treatment of survivorship-induced spurious persistence. The citation
supports the mechanism, not the magnitude. Distinct from the
1995 Brown/Goetzmann/Ross entry at T2. (Section
8)[T3]/[T4] where Section 7 uses it for
the retail options expected-value figure. (Sections 5, 7)[T1] in
Section 7 for the 19,642-trader panel; [T2] in Section 3,
where the byline discrepancy against the 2025 journal version is flagged
and left unresolved. Two sources supply incompatible DOIs (a
Social Science Research prefix vs Brazilian Review of
Finance); neither is carried. (Sections 3, 7,
12)[T1] on the mathematics,
[T6] on the sourcing. (Sections 3, 6)[T1]; the 0.22%-per-standard-deviation point
estimate attributed to it is [T6], unverified
against the paper. (Section 7)min(x, M − x) on 2-of-3 majority; a
competing definition ("wagering the maximum possible stake on every
round") was excluded. (Sections 3, 7, 9, 12)[T2] where
Section 5's Table C uses it for the size factor. (Section
5)[T1] on
the mathematics (the two-barrier gambler's-ruin construction);
[T6] on the sourcing, as the work never
reaches its citing report's own source list. (Section 3)[T6]. (Section 5)[T1] on the mathematics;
[T6] on the sourcing, as the work never
reaches its citing report's own source list. (Section 3)f* = (bp − q)/b adopted; a competing variant
f* = p/(1+g) − q/g excluded as arithmetically wrong. A
sizing rule f* = I(X;Y)/H(X) derived from Kelly's identity
was excluded as a category error. (Sections 3, 6, 7, 9)[T1] on the mathematics,
[T6] on the sourcing. (Sections 3, 6)NobBS method). PLOS Computational Biology 16(4):
e1007735. DOI 10.1371/journal.pcbi.1007735. (Section 6)BS = REL − RES + UNC).
Monthly Weather Review 101(7), 603–608. DOI
10.1175/1520-0493(1973)101<0603:HATMOT>2.0.CO;2. — Section 8
downgrades an uncited invocation of the same
decomposition to [T6], flagging that it appears in no
bibliography there. (Sections 6, 8, 12)[T2] with DOI 10.1016/S0304-405X(99)00022-4 attached to a
Journal of Finance venue label; that prefix mismatch is
flagged and both identifiers are recorded. (Sections 7,
8)[T2] in Section 5's Table C.
(Section 5)[T6] on the authorship of this version. (Section
3)[T2] on the citation (the paper exists and its topic
matches); [T6] on the directional claim
drawn from it — that a time discount can make timid play outperform bold
play — which is unverified in this merge, and whose quantitative
implications for a 90-day horizon the citing source concedes are
uncharacterised in the literature. (Section 3)f* = I(X;Y)/H(X)
attributed to this work was excluded as a category error: Kelly's
identity equates a growth rate with mutual information, whereas
a bet fraction is a dimensionless capital share. (Section
6)[T6] on the
citation trail: cited in one source's body and technique table with no
bibliographic record anywhere in that file. (Section
6)P₊ ≈ 0.30 was
excluded as out of domain. (Section 3)[T6]. Second author's given name (not specified in source).
(Section 6)vollib. (Section 9)[T2]/[T3] in the source. (Section
6)[T2] rather than [T1] because the citing
report's own digest could not confirm the publication details.
(Sections 7, 12)[T5]/[T2]
in the source; a full official-reporter appellate citation, and the
basis on which an unsourced contrary negative was rejected. (Section
11)[T6]/unknown.
The accompanying release number falls in the 92xx-26 series that Section
11 downgraded wholesale to [T6] (see §14.6). URL (not
specified in source). (Section 11)[T3] with a standing verification
flag: the TIR appears in no bibliography in its citing report, and that
report's machine-readable appendix sources the same finding to a CPA
firm's marketing blog ([T4]). (Section
11)[T3] rather than
[T5] because the release reaches this report through a
single source that did not retrieve mass.gov. (Section 11)statsforecast project README
benchmark claim — a 20× speedup over pmdarima and 1.5× over
R's forecast. Vendor benchmark; weight accordingly. URL
(not specified in source). (Section 9)[T6] in
Section 4 for the absence of a docket number and [T4] in
Section 11. Every source describing the injunction describes it as
reaching sports contracts and requiring geofencing of
Massachusetts residents; whether it reaches non-sports economic,
monetary-policy, or election contracts is [T6]/unknown, and
the Massachusetts SJC had not ruled at the research date. (Sections
4, 11, 12)[T4] because the methodology
is contested and not peer-reviewed. (Section 7)https://kalshi.com/docs/kalshi-fee-schedule.pdf. — source
of the taker-fee formula ceil(0.07 × N × P × (1−P)) per
side, adopted 2-of-3 over a derived 1.75% coefficient that its own
reporting source self-tagged [T6]. Whether a
settlement-side fee exists is established by no source, leaving a ~2×
uncertainty on every Kalshi round-trip figure. (Sections 4, 5,
12)[T4] book provenance, not on peer-reviewed methodological
validation, and that the claimed superiority of triple-barrier
over fixed-horizon labeling has not been independently replicated in a
controlled comparison. (Sections 7, 8, 9)[T4] for the claim that estimation error
is the dominant source of long-run Kelly underperformance, and at
[T6] for the far stronger claim that "a 20% relative
misestimation of μ or σ² is enough to make
full Kelly worse than half Kelly for any plausible
parameter vector" — a universal quantifier on a specific threshold,
sourced only to "the MacLean–Ziemba literature" with no named work.
The second claim is unverified in this merge.
(Section 3)https://en.wikipedia.org/wiki/Polymarket.
Downgraded from [T5] for exactly that reason, in a document
that cites three CFTC release numbers elsewhere. (Sections 4,
11)[T4] with the sourcing defect stated; not
upgraded. (Section 5)[T4] with the mismatch stated; not treated as a
Kalshi finding. (Section 5)Grouped into two alphabetical blocks: (a) legal and regulatory authorities, (b) exchange, venue, data-source, and software documentation. A single alphabetical run mixing statutes with package names would be unusable.
[T6]; the three-in-twelve-months figure is
[T4], sourced to investor-education material rather than
rule text. (Sections 4, 7, 11)[T6]. (Sections 5, 11)https://www.ecfr.gov/current/title-31/section-1020.220. —
cited for the written Customer Identification Program and for ACH
funding holds of 3–5 business days (~5.5% of the 90-day clock).
Part 1020 is the banks subpart; see the CIP dispute at
§14.6. (Sections 4, 11)https://www.finra.org/rules-guidance/rulebooks/finra-rules/2090.
— account opening requires SSN or ITIN, address, employment, financial
profile, and risk-tolerance disclosure. (Sections 4, 11).05
supplementary pincite offered for that carve-out is downgraded to
[T6] (see §14.6); the conclusion stands at
[T5]. (Section 11)https://www.finra.org/rules-guidance/rulebooks/finra-rules/2360.
— governs options account approval and firm diligence. The
Levels 1–4 ladder itself is industry convention set firm by firm,
[T4], not a FINRA-codified schedule; a source
presenting it as "codified at 2360(b)(11)–(12)" was resolved against.
Account seasoning of ~30 days for Level 3 and ~60 days for Level 4 is
broker policy at [T6]. (Sections 4, 11)https://www.finra.org/rules-guidance/notices/26-10. —
reported rescission of the pattern-day-trader regime effective June 4,
2026, replaced by intraday margin standards, with broker phase-in
reported through October 20, 2027. Graded [T5] in Section 4
as the better-cited dated claim; downgraded to
[T6]/unsettled in Section 11, which notes the
notice was never retrieved and reaches the merge through a tertiary
chain. Operationally moot either way. (Sections 4, 11)[T6] and
mutually inconsistent across one source's own adjacent table rows; the
concepts (regulated futures contract, qualified board or exchange,
nonequity option) are settled. (Sections 11, 12)[T5]; its
judicial construction is [T6]/unknown, because the
only case authority offered was dropped (see §14.7). (Section
11)[T6]; conformity is
date-limited rather than rolling. (Section 11)[T3] with a standing instruction
to verify against primary text before acting (see TIR 15-14 at T3).
(Sections 11, 12)[T5]; the
effective date is [T6], stated two incompatible
ways within one source ("post-July 4, 2025" throughout, and "the
2018–2025 period" once). Must be verified against the enacted text
before any after-tax calculation. (Section 11)https://www.sec.gov/rules/final/34-96930.pdf. — T+1
settlement. A competing T+2 assertion, made three times by one
source and load-bearing for its turnover analysis, was pruned as
stale; the correction roughly doubles achievable round-trip
frequency in a cash account. Under T+1, unsettled proceeds cannot fund
the next purchase: roughly one round trip per two business days, ~30
round trips across 90 days at best-case timing. (Sections 4, 7,
11)[T5]
in Section 4; [T6]/unsettled in Section
11, unverified against primary source. (Sections 4,
11)https://www.sec.gov/investor/alerts/cashaccounts.pdf. —
cited for the free-riding 90-day cash-up-front restriction. The
three-good-faith-violations trigger count sourced here is
[T4]. (Sections 4, 11)Data-source documentation (Section 10, Table F). PIT
= point-in-time; SBF = survivorship-bias-free. Every
property below is as recorded in Table F.
alfred.stlouisfed.org. — archived vintages of FRED macro
and rates series; API key required. The only genuinely free
source in Table F marked PIT = yes for series with archived
vintages. (Section 10)bankofengland.co.uk/boeapps/database. — UK rates, yield
curves, macro and financial series; PIT = no without vintage storage.
(Section 10)apps.bea.gov/API. — GDP, personal income, trade, industry
accounts; API key; PIT = no without vintage capture. (Section
10)bls.gov/developers. — CPI, employment, wages, productivity,
release calendar; documented daily request and row limits; PIT = no
unless release vintages are stored. Also the free primary source for the
release calendar, which supplies event timing but no historical
consensus forecasts. (Section 10)api.binance.com. — crypto
spot and futures, 2017–present; 1,200 req/min; geofenced for US
IPs, Binance.US required. PIT = partial and SBF = no by
cross-source adjustment (exchange-native sources cover only their own
venue and listed-instrument lifecycle). (Section 10)docs.deribit.com. — BTC/ETH
options, futures, order books, trades, instrument metadata;
exchange-specific published rate limits; public endpoints free but
terms-sensitive. Exchange-specific and not a US equity-options
substitute. (Section 10)data.ecb.europa.eu. —
euro-area rates, FX, macro, banking and financial statistics; PIT = no
unless vintage/release metadata is retained. (Section 10)api.stlouisfed.org. — US and international macro, rates,
labor, prices; 120 req/min; API key. PIT = no on default
endpoints; yes only when realtime_start,
realtime_end, or vintage_dates are
used — the qualifier is preserved because dropping it licenses
exactly the error the column exists to prevent. (Section
10)gdeltproject.org. — global news
events, tone/sentiment, entity extraction; multi-decade corpus; openly
accessible. Not a licensed article-text feed; known
defects include source duplication, language imbalance, timestamp
ambiguity, and entity-resolution errors. (Section 10)docs.kalshi.com. —
CFTC-regulated event contracts; order books, trades, settlements, market
status and history; 2021–present. Basic tier runs a token
bucket: 200 read + 100 write tokens/sec against a default request cost
of 10 tokens, i.e. ~20 read and ~10 write req/sec; 429 responses omit
Retry-After, so client backoff must be
self-managed. PIT = no, SBF = no — market-history retention is
endpoint-specific with no blanket guarantee, and the remedy is local
archiving at capture time. Automated access is supported;
redistribution and derived commercial use remain restricted.
No first-class Python client existed on PyPI as of
2026-08-01 — a competing assertion of
kalshi-python-async was excluded, so the client must be
hand-rolled. (Sections 9, 10)api.kraken.com. — crypto spot,
inception–present; 1 req/sec public, REST OHLC capped at 720 bars per
request. PIT = partial, SBF = no by the same exchange-native adjustment
applied to Binance. (Section 10)data.sec.gov. — US filings,
XBRL company facts, submissions, full-text search; filings ~1990s
onward; real time on filing acceptance; 10 req/sec fair-access ceiling;
descriptive User-Agent required, no key. PIT =
partial — reconstructable if indexed by filing acceptance
timestamp, but the raw facts API is not itself a PIT database, since
amendments and taxonomy changes require event-time filtering.
SBF = partial — filings of delisted issuers are
retained, but EDGAR supplies no priced security master,
so a survivorship-bias-free equity universe cannot be built from it. A
competing yes/yes grading was resolved against, and an internal
contradiction in another source ("survivorship-bias-free: yes" alongside
a concession that delisting-aware prices require a paid vendor) was
corrected. (Sections 8, 10)fiscaldata.treasury.gov; home.treasury.gov. —
par yield curves, bill rates, auctions, debt, receipts and outlays; no
universal documented quota; PIT = no for revised series. (Section
10)query1/query2.finance.yahoo.com) and the
yfinance wrapper. — US and global equities, ETFs, FX,
crypto, options chains, fundamentals, news; ~30y daily, ~60d intraday,
provider guarantees no retention. No published
rate limit; ~2,000 req/hr is a community-observed throttling
threshold, labelled unofficial by the source that supplies it.
PIT = no, SBF = no, and terms restrict automated extraction and
redistribution — the only source in one report's entire audit
with tos_restricts_automation: true. Cookie/crumb handshake
behaviour and endpoint schemas have already changed; IP bans reported.
Both surviving sources converge independently on restriction; neither
quotes the governing clause, so this is a legal-review item, not a
settled fact. (Sections 8, 9, 10)Python package documentation (Section 9, Table E). Entries carry
[T5] where both surviving sources agree on the version, or
where a licence or API claim was corroborated to a primary project
document. All star and open-issue counts across Table E are
[T6] without exception — neither source evidenced a single
API response. Versions should be re-pinned at install
time.
arch 8.0.0 (2025-10-21). Licence
NCSA. pypi.org/pypi/arch/json ·
github.com/bashtage/arch. — GARCH-family volatility,
unit-root testing, bootstrap inference. Licence resolved to NCSA on an
explicit pyproject.toml declaration, against a competing
MIT claim; version [T6] (8.0.0 vs 7.2.0 across sources).
Provides first-class SPA, StepM, and
MCS classes in arch.bootstrap, plus
StationaryBootstrap — correcting a source claim
that Hansen's SPA test is "not in any single first-party library."
(Sections 8, 9)backtesting (backtesting.py) 0.6.6
(2026-07-22). Licence AGPL-3.0-or-later.
pypi.org/pypi/backtesting/json ·
github.com/kernc/backtesting.py. — event-driven,
single-instrument research loop, single maintainer. Genuine copyleft:
distributing a derivative, including over a network, triggers
source-disclosure obligations. (Section 9)cmdstanpy 1.3.0 (2025-10-20). Licence
BSD-3-Clause. pypi.org/pypi/cmdstanpy/json ·
github.com/stan-dev/cmdstanpy. — Stan HMC/NUTS interface;
two-source version agreement. (Sections 6, 9)cvxpy 1.9.2 (2026-06-22). Licence
Apache-2.0. pypi.org/pypi/cvxpy/json ·
github.com/cvxpy/cvxpy. — the convex-programming DSL
underneath both PyPortfolioOpt and
Riskfolio-Lib; two-source version agreement. (Sections
6, 9)ib-async 2.1.0 (2025-12-08). Licence
BSD-2-Clause [T6]. pypi.org/pypi/ib-async/json
· github.com/ib-api-reloaded/ib_async. — the only actively
maintained Python framework speaking Interactive Brokers' native
protocol. Repository resolved to ib-api-reloaded against a
competing erdewit URL, which is the pre-transfer location:
the package was renamed and transferred after the original maintainer's
death in early 2024, superseding the abandoned ib_insync.
(Section 9)linearmodels 7.0 (2025-10-21). Licence
NCSA. pypi.org/pypi/linearmodels/json ·
github.com/bashtage/linearmodels. — panel data,
instrumental variables, asset-pricing factor models. Same maintainer and
licence-declaration mechanism as arch. (Section
9)mlflow 3.15.0 (2026-07-31). Licence
Apache-2.0. pypi.org/pypi/mlflow/json ·
github.com/mlflow/mlflow. — experiment tracking; two-source
version agreement. The local filesystem backend suffices at this scale.
(Section 9)nautilus-trader 1.230.0 (2026-06-29).
Licence LGPL-3.0-or-later.
pypi.org/pypi/nautilus-trader/json ·
github.com/nautechsystems/nautilus_trader. — Rust core,
true event-driven order semantics, live and backtest from the same code
path; steepest learning curve in the survey. Version resolved to the
higher figure against a competing claim pairing a lower version
with a later date. (Section 9)numpyro 0.21.0 (2026-05-02). Licence
Apache-2.0. pypi.org/pypi/numpyro/json ·
github.com/pyro-ppl/numpyro. — JAX-backed probabilistic
programming; two-source version agreement. (Section 9)Riskfolio-Lib 7.3.0 (2026-05-31).
Licence BSD-3-Clause. pypi.org/pypi/riskfolio-lib/json ·
github.com/dcajasn/Riskfolio-Lib. — 26 convex risk
measures, risk-parity variants, hierarchical clustering,
Black-Litterman, entropy pooling. A fully-qualified Kelly API
path (Riskfolio-Lib.optimization.mean_risk.portfolio_kelly)
offered by one source is [T6] and explicitly not to be
trusted — stated with precision, zero documentation links, and
no version pin. (Section 9)scikit-learn 1.9.0 (2026-06-02).
Licence BSD-3-Clause. pypi.org/pypi/scikit-learn/json ·
github.com/scikit-learn/scikit-learn. — the
strongest metadata corroboration anywhere in Section 9: both
sources agree on version and release date to within one day. Supplies
TimeSeriesSplit; does not implement purged or
embargoed cross-validation. (Sections 8, 9)skfolio 0.20.1 (2026-04-21). Licence
BSD-3-Clause. pypi.org/pypi/skfolio/json ·
github.com/skfolio/skfolio. —
skfolio.model_selection.CombinatorialPurgedCV and
.WalkForward are the single most useful finding in Section
9: a maintained, BSD-3-licensed, sklearn-API-compatible
implementation of combinatorial purged cross-validation, in a package
already needed for portfolio optimization. The best-sourced API claim in
that cluster, naming the package's __init__ exports.
(Sections 8, 9)sktime 1.1.0 (2026-07-28). Licence
BSD-3-Clause. pypi.org/pypi/sktime/json ·
github.com/sktime/sktime. — unified forecasting API over
500+ models; two-source version agreement. (Section 9)statsmodels 0.14.6 (2025-12-05).
Licence BSD-3-Clause. pypi.org/pypi/statsmodels/json ·
github.com/statsmodels/statsmodels. — econometrics plus
statsmodels.tsa.statespace for SARIMAX,
UnobservedComponents, VARMAX, and Kalman filtering. Version agreed
across sources; release date [T6] (two incompatible dates
for one version). (Sections 6, 8, 9)vectorbt 1.1.0 (2026-07-05). Licence
Apache-2.0 with the Commons Clause addendum.
pypi.org/pypi/vectorbt/json ·
github.com/polakowo/vectorbt. — vectorized, Numba-compiled
parameter sweeps. Licence resolved to include the Commons Clause against
a bare Apache-2.0 claim: the specific, falsifiable claim beats the SPDX
field GitHub returns, which does not represent addenda. Free for
individuals; bars selling a product whose value derives primarily from
the software. (Section 9)Entries here are retained as citations because a section presented them as a source, and are labelled as unreliable because that section's resolution said so. None should be relied on without independent verification.
Legal, regulatory, and tertiary sources
[T1] by one source, made load-bearing for its entire
Kelly-inadequacy argument, then omitted from its own 47-entry
bibliography. No source in this merge supplies a venue, page
range, or DOI, and none is invented. (Section 3)[T5] in Section 4.
(Sections 4, 11)p₀, and it
generates neither the one-sided nor the two-sided figures in that table.
The formula and its accompanying claim that "the two methods
agree to within rounding" were dropped; the citation is
recorded here so it is not reintroduced. (Section 8)[T5] without it. (Section 11)[T4] book provenance, not on peer-reviewed
methodological validation, and this report does not claim otherwise.
(Section 8)[T5] and corroborated by
both sources. (Section 11)https://en.wikipedia.org/wiki/Kalshi. — the only
source cited by any report for the Commonwealth v.
KalshiEX preliminary injunction and the Massachusetts
sports-contract geofencing requirement, in Section 4's structural-gates
table. (Sections 4, 11)https://en.wikipedia.org/wiki/Polymarket. — the
only source cited by any report for the Polymarket federal
chronology: the QCEX acquisition, the December 2, 2025 US unblocking
date, the July 15, 2025 closure of the DOJ/CFTC investigations, and the
November 2025 Amended Order of Designation. One source tags all six
chronology items [T5] — its own tier for primary regulatory
documentation — while sourcing every one of them to this tertiary page.
(Sections 4, 11)[T6] with "primary citation withheld
pending verification", and neither sibling report mentions the
instrument. No feasibility verdict in this report rests on
it. (Section 4)Python package documentation graded [T6] in
Section 9 — single-source assertions, version conflicts, or unverified
metadata
Every entry below carries
pypi.org/pypi/<name>/json plus the stated GitHub
repository as its consulted source URL, exactly as recorded in Table E.
All star and open-issue counts are [T6] without
exception; neither surviving source captured a single API response, no
HTTP status codes, no retrieval timestamps, and no ETags.
Verify versions at install time.
arviz 1.2.0 (2026-06-12), Apache-2.0 —
github.com/arviz-devs/arviz. Posterior diagnostics: ESS, R̂,
LOO, WAIC. A Bayesian forecast published without R̂ and ESS is an
unaudited number. (Sections 6, 9)bt 1.2.0 (2026-04-25), MIT —
github.com/pmorissette/bt. Tree-structured portfolio
composition. (Section 9)bayesian-changepoint-detection —
version and repository (not specified in source). Existence
unverified. (Section 6)catboost 1.2.8 (2026-04-29),
Apache-2.0 — github.com/catboost/catboost. Single-source
(one report omitted the gradient-boosting tier entirely from 53 rows).
(Section 9)ccxt 4.5.70 (2026-07-29), MIT —
github.com/ccxt/ccxt. Unifies 105+ crypto venues. CCXT Pro
(WebSocket streaming) is a paid product; the open-source package is
REST-only. (Section 9)chainladder — version and repository
(not specified in source). Loss development. Relevant only to
the one technique Section 6 explicitly disclaims as
non-transferable. (Section 6)dagster 1.13.16 (2026-07-30),
Apache-2.0 — github.com/dagster-io/dagster. Asset-centric
DAG orchestration; the alternative to prefect. Pick one.
(Section 9)darts 0.46.1 (2026-07-20), Apache-2.0
— github.com/unit8co/darts. Version conflict across sources
(0.46.1 vs 0.31.0); a 15-minor-version gap is implausible as noise.
(Section 9)duckdb 1.5.5 (2026-07-22), MIT —
github.com/duckdb/duckdb. Analytical SQL over partitioned
Parquet; no server, no cloud dependency, no cost. Star count resolved
against the main project repo rather than the duckdb-python
sub-repo. (Section 9)FinancePy 1.0.1 (2025-08-31),
GPL-3.0-or-later —
github.com/domokane/FinancePy. Slowing cadence; overkill
for a pure-options track. Author's name misspelled in one source.
(Section 9)gluonts 0.17.0 (2026-07-31),
Apache-2.0 — github.com/awslabs/gluonts. Reference
implementation of DeepAR; cadence slowing. (Section 9)hierarchicalforecast 1.5.1
(2026-03-04), Apache-2.0 —
github.com/Nixtla/hierarchicalforecast. The only
library in the surveyed universe addressing hierarchical
reconciliation. Single-source, retained because it answers a
required capability nothing else covers. (Section 9)lifelib — version and repository (not
specified in source). Actuarial projection primitives. (Section
6)lightgbm 4.7.0 (2026-05-04), MIT —
github.com/microsoft/LightGBM. Single-source. Wrap in
skfolio's CV splitters, never in
sklearn.model_selection.KFold. (Section 9)neuralforecast 3.2.0 (2026-07-10),
Apache-2.0 — github.com/Nixtla/neuralforecast. Included for
forward extension; not needed for the doubling experiment. (Section
9)numpy 2.5.1 (2026-07-04), BSD-3-Clause
— github.com/numpy/numpy. Substrate. (Section
9)optuna 4.9.0 (2026-06-01), MIT —
github.com/optuna/optuna. Hyperparameter search.
Use sparingly and inside purged CV: an HPO loop over a
short financial series is a machine for manufacturing overfit Sharpe
ratios, and Section 8's multiple-testing corrections apply to every
trial it runs. (Section 9)pandas 3.0.5 (2026-07-22),
BSD-3-Clause — github.com/pandas-dev/pandas. Substrate.
(Section 9)pandas-datareader 0.11.1 (2026-06-24),
BSD-3-Clause — github.com/pydata/pandas-datareader. FRED,
World Bank, OECD. Marked "active" on a 2026-06-24 release while
elsewhere described as having "the first release in over a year" —
re-verify before depending on it. Its Yahoo path broke in 2020 and has
not returned. (Section 9)pandera 0.32.1 (2026-06-29), MIT —
github.com/unionai-oss/pandera. DataFrame and Series schema
validation. Gate every ingested DataFrame behind a schema before
it reaches the analytical pipeline — a check that fires on a
silently changed yfinance column layout is worth more than
any downstream defensive coding. (Section 9)pmdarima 2.1.1 (2025-11-17), MIT —
github.com/alkaline-ml/pmdarima. Slowing —
one release in the trailing 365 days, lead maintainer redirected to
statsforecast. Retained as a reference implementation;
prefer statsforecast.AutoARIMA for new code. (Section
9)polars 1.43.2 (2026-08-01), MIT —
github.com/pola-rs/polars. Substrate and feature
computation. Caveat: the stated release date is the same
calendar day as the research date — a zero-day-old release is exactly
the shape of a value generated to satisfy a recency rule rather than
observed. (Section 9)polygon-api-client 1.16.3
(2025-10-30), MIT — github.com/polygon-io/client-python. US
equities and options chains. A claimed successor package
(massive 2.8.0, following an asserted Polygon→Massive
rebrand) was excluded: the asserted rebrand date is
identical to this package's asserted release date, a coincidence the
citing report's own digest flags. (Section 9)prefect 3.8.1 (2026-07-30), Apache-2.0
— github.com/PrefectHQ/prefect. Orchestration; both sources
independently select it for the reference architecture. (Section
9)purgedcv ≥ 0.1.3 (PyPI,
scikit-learn-compatible). Repository (not specified in source). —
recommended by one source, whose own text dates the release to the same
day as its report, describes it as a mature mlfinlab
replacement, and elsewhere concedes it does not expose the functions it
is recommended for. Verify independently before
adopting. (Section 8)pyarrow 25.0.0 (2026-07-10),
Apache-2.0 — github.com/apache/arrow. Zero-copy columnar
interchange. (Section 9)pydantic 2.13.4 (2026-05-06), MIT —
github.com/pydantic/pydantic. Record, config, and contract
validation. Version conflict; the competing claim pairs a lower version
with a later date. (Section 9)pyfeng 0.5.0 (2026-05-26),
GPL-2.0 — github.com/PyFE/PyFENG. SABR,
Heston, NSVh, Schöbel-Zhu, rough Heston. The strictest copyleft in the
stack; appropriate for private research. (Section 9)pyirt — version and repository (not
specified in source). Item-response theory. statsmodels
does not ship IRT. (Sections 6, 9)pykalman 0.11.2 (2026-01-31),
BSD-3-Clause — github.com/pykalman/pykalman. Optional;
statsmodels.tsa.statespace covers the same ground with
better integration. (Section 9)pymc 6.2.0 (2026-07-23), Apache-2.0 —
github.com/pymc-devs/pymc. Major-version conflict
across sources (6.2.0 vs 5.17.0) — resolve at pip install
time. If 6.x is real, expect breaking API changes against every
PyMC tutorial written before it. (Sections 6, 9)PyPortfolioOpt 1.6.0 (2026-02-26), MIT
— github.com/robertmartin8/PyPortfolioOpt. Mean-variance
and Black-Litterman, slowing cadence. (Section 9)pytensor 3.2.3 (2026-07-25),
BSD-3-Clause — github.com/pymc-devs/pytensor. The
Theano/Aesara successor underneath PyMC. (Section 9)QuantLib (Python bindings) 1.43
(2026-07-14), BSD-3-Clause (QuantLib modified) —
github.com/lballabio/QuantLib-SWIG. Version and repository
conflicted; lballabio is upstream, not a
mirror — one source inverted the canonical/mirror relationship.
The only repository-attribution conflict resolved in the competing
source's favour. (Section 9)ruptures — version conflict (v1.1.9 vs
"1.x range"); repository (not specified in source). Change-point
detection. A cited API path, ruptures.detect.cusum,
does not exist: the package exposes search classes
Pelt, Binseg, Window,
BottomUp, Dynp with cost functions.
(Sections 6, 8, 9)scipy 1.18.0 (2026-06-19),
BSD-3-Clause — github.com/scipy/scipy. Substrate. A
cited symbol, betaind "from scipy," does not
exist; the correct symbol is scipy.stats.beta.
(Sections 6, 8, 9)scores 2.6.0 (2026-07-17), Apache-2.0
— github.com/nci/scores. Brier and threshold-Brier, CRPS,
FIRM, SEEPS, Kling-Gupta Efficiency, the Diebold-Mariano test, Fractions
Skill Score, isotonic regression for reliability diagrams.
Repository resolved to nci/scores against a
competing owner flagged as fabricated-looking by its own
digest. Caveat: a worked recipe calls
scores.brier_score and scores.crps_ensemble as
top-level functions, but the package organizes metrics under submodules
— read the module layout before writing imports. (Sections 6,
9)scoringrules 0.11.0 (2026-06-06),
Apache-2.0 — github.com/frazane/scoringrules. CRPS, energy,
variogram, interval, and quantile scores across NumPy, JAX, PyTorch,
TensorFlow backends. (Section 9)statsforecast 2.1.1 (2026-07-16),
Apache-2.0 — github.com/Nixtla/statsforecast.
Numba-compiled AutoARIMA, AutoETS, MSTL, Theta, CES. Version conflict
resolved to the higher figure. (Section 9)vollib 1.0.11 (2026-06-01), MIT —
github.com/vollib/py_vollib. Black, Black-Scholes,
Black-Scholes-Merton prices; the full standard Greek set; implied
volatility via Jäckel's Let's Be Rational. Use
vollib, not py_vollib, which is a
deprecated transitional alias — though note the anomaly that the
deprecated shim carries version 1.0.12 against the canonical package's
1.0.11, which is backwards, with both given identical release dates. At
least one of those version numbers is wrong. (Section 9)xgboost 3.3.0 (2026-06-20), Apache-2.0
— github.com/dmlc/xgboost. Single-source. (Section
9)xskillscore 0.0.29 (2026-02-18),
Apache-2.0 — github.com/xarray-contrib/xskillscore. The
right tool only if forecasts live in xarray. (Section 9)yfinance 1.5.2 (2026-07-23),
Apache-2.0 — github.com/ranaroussi/yfinance. Default for US
equity and ETF end-of-day OHLCV, dividends, splits. Its
constraint is legal, not technical — see the Yahoo Finance
entry at T5(b). The most-used ingestion package in retail quant is the
one operating furthest outside its provider's terms, and it is flagged
PIT = no, SBF = no, ToS-restricted. (Sections 8, 9, 10)These were dropped as fabricated, unlocatable, misattributed, or not supporting the proposition attached to them. They are listed here solely so that a downstream reader, editor, or model does not reintroduce them. None is a citation of this report.
Struck as fabricated or non-existent
kalshi-python-async 3.25.0 — asserted with a star count
and a repository URL falling in the set flagged as fabricated-looking;
an affirmative PyPI search found no Kalshi Python client. (Section
9)massive 2.8.0; quantstats (recommended
twice with no version, date, licence, repository, or star count, in a
report built entirely on per-package metadata). (Section
9)[T6] and not to be
quoted as a peer-reviewed result. (Section 8)10.1214/aos/1176348899 and
10.1093/rfs/4.4.867 — absent from their own bibliography
and mismatched to the methods named. (Section 8)[T1] grade. (Section
7)10.1111/0022-1082.00341. (Section 3)Struck as real works that do not support the proposition attached
[T6]/unknown. (Section
11)[T1] — neither
URL supports the graded claim. (Section 5)[T6]). (Section 7)[T1] citation replaced by Mitchell
& Pulvino; the foundational volatility-risk-premium and
"Volatility-of-Volatility Risk" citations replaced by Carr & Wu and
Coval & Shumway; the PEAD decay citation (real paper, wrong topic);
the momentum net-Sharpe source; the index-reconstitution price-impact
citation (wrong authors); and the low-volatility citation (confabulated
authorship, row retained with no citation asserted and none
invented). (Section 5)| Tier | Entries |
|---|---|
| T1 | 119 |
| T2 | 50 |
| T3 | 12 |
| T4 | 19 |
| T5 | 99 — of which 56 legal/regulatory authorities and 43 documentation sources (28 data sources, 15 Python packages) |
| T6 | 72 — of which 29 legal, regulatory, tertiary or unnamed sources and 43 Python packages |
| Total distinct citations | 371 |
Incomplete bibliographic information. 157 entries carry at least one field marked "(not specified in source)" — counted per entry, not per field, so an entry missing both a year and a DOI is counted once. That is 42% of the bibliography, and it is a finding about the source corpus rather than an artifact of compilation: the sections went to considerable length to refuse DOIs that were disputed, venue-mismatched, or absent, and those refusals are reproduced here rather than repaired.
Separately, and overlapping with the above:
40 struck citations are registered in §14.7 and are deliberately absent from the bibliography proper.
This section argues against the report. It is not a disclaimer, and it does not exist to make the preceding twelve sections feel more rigorous by gesturing at humility. Every item below is a specific, checkable way the merged findings could be wrong, ordered by how much damage each would do if it were.
Start with the structural fact that governs all of it: this report is a merge of three AI-generated deep-research outputs, adjudicated largely by majority rule, with no human literature review anywhere in the chain. Nothing downstream of that fact is stronger than that fact.
The report's tier tags create an appearance of uniform grading that
the underlying evidence does not support. Four sections — 5, 6, 7, and 8
— each independently state that they re-derived tier tags rather than
inheriting them, and each states that its counts are not comparable to
any source report's counts. That means the report's own quality
instrument is not calibrated across its own length. A [T5]
in Section 4 and a [T5] in Section 11 were assigned by
different synthesizers against the same specification with no
cross-checking. Readers who compare tier density across sections are
comparing nothing.
Beneath that, the coverage is badly uneven:
[T3], and TIR 15-14 does not appear in Gemini's
own bibliography while its JSON sources the same finding to a CPA firm's
marketing blog.P₊, no P₀, no τ₊ distribution for
this objective exists in any of the three reports or in any literature
they identify. The 1%–5% base rate is an extrapolation across
populations, horizons, and capital scales that no study measures
directly.The Dubins–Savage anchor does not cover this problem, and the
report says so before reasoning from it anyway. Section 3.4
states explicitly that bold-play optimality is a result about the
unbounded-time goal problem, that the finite-horizon optimum is
the solution to an HJB equation with terminal condition
V(w,T) = 1{w ≥ 200}, that this is not in general bold play,
and that no source in the merge solves it. Sections 7.3 and 12.2 then
both reason from bold play. If Chen's (1977) discount-factor inversion
holds directionally — MiniMax's own open-questions section concedes the
quantitative implications at a 90-day horizon are uncharacterised — then
timid play can dominate, and the report's central prescriptive intuition
(concentration buys upper-tail mass, diversification destroys it) loses
its theoretical anchor for the objective actually posed.
The binary utility is almost certainly not the principal's
utility. Section 3.3 derives its whole bold-play direction from
U(W_T) = 1{W_T ≥ 200} — a utility indifferent between USD
100 and USD 0. Real principals are not. The moment the retained stake
carries value, V = 100·(P₊ − P₀) stops being the right
value function, and Table G's ordering by P(reach) stops being the right
ordering. The report never tests its ranking against any other
utility.
Table G ranks on one criterion and excludes on another. Section 12.2 ranks by P(reach) descending. Section 12.3 excludes leveraged and inverse ETFs on expected-value grounds. But Section 7.3's most interesting correction is precisely that leverage raises P(reach) while lowering EV, and that conflating the two "collapses the entire first-passage framing the report is built on." The exclusion table performs the conflation Section 7.3 diagnosed. A 2× ETF needs +41% on the underlying, not +100%; on the report's own stated ranking criterion it has a claim to a row.
The P(ruin) headline measures a threshold the merge chose. Section 12.1 redefines ruin from MiniMax's stated "≤ USD 0" to "≤ USD 25," on the grounds that the original made roughly half the column impossible. That is the right repair, but it means the widely-quotable "60% probability of an experiment-killing loss" is a function of a cutoff invented during the merge. At a USD 10 floor the number falls; at USD 50 it rises. No source supports any particular value.
Package and vendor risk is under-modelled. Section 9
designates skfolio.model_selection.CombinatorialPurgedCV as
"the single most useful finding in the cluster" — a maintained,
sklearn-compatible purged cross-validator. It is single-sourced to
MiniMax, the same source that invented
ruptures.detect.cusum and betaind "from scipy"
and a fully-qualified Riskfolio-Lib Kelly path its own
digest called "the least sourced" line in the section. The import is
better-evidenced than those (Section 9 names the package's exports and
shows the statement), so this is a flagged risk rather than an
accusation. But if it is wrong, the report has no
answer to its own purged-CV requirement: mlfinlab is
commercial and never on PyPI, mlfinpy is dead at 661 days,
timeseriescv is seven years dormant. Similarly, Section
10's entire event-contract data path assumes Kalshi's terms continue to
permit automated access; a terms change makes Section 12's row 3
unbuildable even if the Massachusetts question resolves favorably.
The friction numbers are the least-measured numbers in the report, and they are the ones that produce its conclusion. Section 4 states in its own header that no cell in Table A is a point estimate, that MiniMax conceded its estimates "may understate actual retail friction by 50%–200%," and that no live spread tape was pulled for any pair by any source. Sections 5, 7, and 12 then treat those cells as the binding constraint that defeats every documented edge. The numbers that produce the headline are the numbers the report itself declined to measure.
All three source reports were AI-generated. That has three specific consequences, and none of them is speculative here — each is documented inside this report.
Fabrication at scale. Independent verification found roughly one third of MiniMax's Cluster 3 citations fabricated or materially misattributed, including a paper attributed to "Penn (2025)" that does not exist and whose surname matches this project's commissioner. Gemini's self-declared "Chain-of-Verification Audit" is disproven from inside its own deliverable. Qwen attributed post-earnings drift to a paper about IPO underperformance, short-term reversal to a momentum paper, cited an Instagram Reel as the source for Wald (1945), and an ACL NLP findings paper for Brier (1950).
Retrieval bias toward the heavily indexed. Kelly (1956), Fama–French, Barber–Odean, Harvey–Liu–Zhu, Hou–Xue–Zhang appear in all three reports. Anything behind a subscription, in a poorly-indexed venue, in a conference proceeding, or not in English appears in none. The corpus contains zero non-English literature and zero unpublished practitioner data. Meanwhile Wikipedia is the sole cited source for every load-bearing 2025–2026 regulatory fact: the Polymarket chronology and QCEX acquisition, the CFTC Amended Order of Designation, the Commonwealth v. KalshiEX injunction (no docket number from any source), and the reported June 2026 PDT rescission.
Named gaps the merge itself flagged and could not fill. Novy-Marx & Velikov's taxonomy of anomaly trading costs — conspicuous, since friction is the binding constraint everywhere in this report. Chen & Zimmermann's publication-bias work, the standard modern counterweight in the factor-zoo debate. CBOE's free volatility indices and term structure — conspicuous, since both surviving sources declare options data the largest gap in the free universe while neither examines the exchange that publishes it. Finnhub, Databento, Alpaca, OpenBB, Coinbase's public REST API, CFTC Commitments of Traders. The NFA and the SIPC / CFTC Part 190 customer-protection regimes, named by no source, and directly relevant to what happens to the balance if a venue fails. And no constructive options-strategy treatment exists in any of the three reports, which makes Section 5's options coverage unavoidably one-sided.
Temporal coverage is bimodal rather than continuous. Technical analysis rests on 1992 / 2007 / 2012 and nothing after. PEAD rests on 1968 / 1989 / 1990 plus an unnamed working paper. Retail options rests on 2022–2023. The middle is missing.
A quant researcher attacks Section 12's probability column,
and the attack lands. Section 3 formalizes the objective
correctly as a first-passage problem and supplies the finite-horizon CDF
in closed form:
P(τ_B ≤ T) = Φ((θT − B)/(σ√T)) + e^(2θB/σ²)·Φ((−θT − B)/(σ√T)).
The parameters are observable — BTC realized volatility is public,
B = ln 2, T = 63 trading days. The report had
the formula and the inputs and did not evaluate it. Instead, Table G's
entire P(reach) column is tagged [T6] in every cell,
populated from three AI reports' unmodeled assertions, and the
median-days-to-target and IQR columns are null throughout
because no source ever derived a first-passage time
distribution. This is the one attack the report cannot answer. It built
the right instrument in Section 3 and declined to use it in Section
12.
A securities attorney attacks the Massachusetts chain. Not one Massachusetts authority was read. The governing chapter was disputed and resolved by inference. Two Technical Information Releases carry load-bearing findings and neither was retrieved. The injunction that drives the report's operative recommendation has no docket number, no division, no judge, and no reporter from any source. The report reaches a specific, actionable recommendation — do not assume lawful access to non-sports event contracts as a Massachusetts resident — on a record where the primary sources were unreachable.
A CFTC compliance officer attacks the release
numbers. MiniMax cites seven releases in the 92xx-26 series,
none verified against primary source, and one (9276-26) appears in its
bibliography supporting nothing in its body — a hallucination signature.
Release 9240-26 is the sole bridge from BTCPERP to Section 1256
treatment. Section 11 downgrades the entire series to [T6],
which is correct, but Section 12's tax-wedge arithmetic still runs. That
officer would also note that "no CFTC-specific retail empirical
literature exists" may describe where the reports looked rather than
what exists.
This is structural to how the merge was produced and it deserves to be stated without softening.
Three AI systems drawing on overlapping training data and similar retrieval strategies can agree with one another while all being wrong in the same direction. Majority rule detects disagreement; it is blind to correlated error by construction. Where all three reports agree, this merge assigned its highest confidence — but three-way agreement on a heavily-indexed fact is nearly free, and three-way agreement on a rare or contested fact does not occur anywhere in this corpus. The merge caught the mechanism once, in Section 10: MiniMax and Gemini's identical Kalshi rate limits "plausibly read the same documentation page, so this is corroboration rather than independent confirmation." That caveat generalizes to every canonical citation appearing identically in all three reports, and it was applied in one place.
Two further mechanisms compound it:
The resolution rule is inverted against the observed failure mode. Section 11 states its rule as "prefer the more specifically cited claim over the more confidently stated one." Sections 4, 5, and 8 apply variants. But one generator's demonstrated failure mode is specific-looking fabrication — precise DOIs attached to nonexistent papers, fully-qualified API paths to functions that do not exist, seven numbered CFTC releases none of which verify. Against a generator that fabricates specifics, specificity is a weak signal, and a rule that rewards it selects for the fabrication.
Self-audit by the same model family is not independent verification. Throughout the merge, conflicts were adjudicated using each source report's own digest flags — "MiniMax's own digest identifies this as a category error," "flagged by its own audit." Those digests were themselves AI-produced from the same corpus. Where a digest failed to flag something, the merge inherited the miss silently, and there is no way to measure how often that happened.
Section 12.4 reports
positive_expected_value_exists: false and calls it "the
only unanimous finding across all three source reports." Four
objections:
Unanimity is partly an artifact of the prompt. All three reports ran the same CASINO brief, which specifies a Narrator who "refuses to be encouraging." Three compliant models converging on skepticism is weaker evidence than three independent analysts converging on it.
It is a universal negative over an admittedly incomplete survey. The report's own gap list — Novy-Marx & Velikov, Chen & Zimmermann, CBOE, COT, no constructive options treatment, no non-English literature, no CFTC-specific empirics — means "no positive-EV approach exists in the surveyed universe" is a statement about the survey.
Unverified EV is not negative EV, and the report reports the
two as though they were the same. Gemini's Kalshi
favorite-buying strategy at P ≥ 0.70 with maker limit
orders was rejected on three grounds: single-sourced, 2–4× the racetrack
figure, and dependent on an uncited "maker fee = 0%" assumption. All
three are provenance objections. None is evidence that the
strategy's sign is negative. Whether Kalshi charges maker fees
is a checkable fact that nobody checked, and Section 4's own corrected
arithmetic puts the taker fee at P = 0.90 at 1.4% round trip —
a maker pays less. If the maker fee is zero and the favorite-longshot
bias in a retail-dominated venue is even half the racetrack magnitude,
gross EV is positive.
The report's own text contains a positive-EV strategy. Section 5.4.4 calls the volatility risk premium "the most economically robust effect in this section" and excludes it on capital grounds — a USD 2,000 margin floor — not on expected-value grounds. "No positive EV at USD 100 scale" and "no positive EV" are different claims. Section 12's headline states the second.
[T6] guesses with null
time-to-target throughout. This is the single weakest point in the
merged report.{
"research_date": "2026-08-01",
"objective": {
"starting_capital_usd": 100,
"target_capital_usd": 200,
"horizon_days": 90,
"jurisdiction": "US-MA"
},
"vehicles": [
{
"name": "Fractional equity / ETF - liquid large-cap",
"asset_class": "Equities / ETFs",
"min_viable_position_usd": 1.0,
"min_viable_position_pct_of_stake": 1.0,
"round_trip_friction_pct": 0.02,
"feasible_at_100usd": true,
"regulatory_gates": [
"FINRA Rule 4210(b)(4) $2,000 margin minimum forces a cash account",
"SEC Rule 15c6-1 T+1 settlement (eff. 2024-05-28)",
"Regulation T free-riding / good-faith violation, 12 CFR Part 220 (90-day cash-up-front restriction)",
"FinCEN CIP / FINRA Rule 2090 KYC; 3-5 business day ACH funding hold"
],
"venues": [
"Robinhood",
"Fidelity",
"Charles Schwab",
"Interactive Brokers Lite"
],
"notes": "Total round-trip friction band 0.02%-0.30% of stake; low end carried per the stated convention. Unanimous across all three source reports and the only vehicle whose round-trip friction is reliably below 1% of stake. Constraint is not cost but the absence of any structural mechanism to double: 1:1 exposure, no leverage without borrowing, T+1 caps turnover at ~30 round trips over the window."
},
{
"name": "Fractional equity - low-priced / small-cap",
"asset_class": "Equities",
"min_viable_position_usd": 1.0,
"min_viable_position_pct_of_stake": 1.0,
"round_trip_friction_pct": 0.6,
"feasible_at_100usd": true,
"regulatory_gates": [
"FINRA Rule 4210(b)(4) $2,000 margin minimum forces a cash account",
"SEC Rule 15c6-1 T+1 settlement (eff. 2024-05-28)",
"Regulation T free-riding / good-faith violation, 12 CFR Part 220 (90-day cash-up-front restriction)",
"FinCEN CIP / FINRA Rule 2090 KYC; 3-5 business day ACH funding hold"
],
"venues": [
"Robinhood",
"Fidelity",
"Charles Schwab",
"Interactive Brokers Lite"
],
"notes": "Friction band 0.60%-2.10% of stake; spread-dominated (0.50%-2.00% typical spread, [T4]). Table A verdict: 'Yes (spread-dominated)'."
},
{
"name": "Listed options - long single leg",
"asset_class": "Listed equity/index options",
"min_viable_position_usd": 5.0,
"min_viable_position_pct_of_stake": 5.0,
"round_trip_friction_pct": 2.5,
"feasible_at_100usd": false,
"regulatory_gates": [
"FINRA Rule 4210(b)(4) $2,000 margin minimum forces a cash account",
"FINRA Rule 2360 options approval - Level 1-2 only (long calls/puts)",
"SEC Rule 15c6-1 T+1 settlement (eff. 2024-05-28)",
"Regulation T free-riding / good-faith violation, 12 CFR Part 220 (90-day cash-up-front restriction)"
],
"venues": [
"Robinhood",
"Charles Schwab",
"Interactive Brokers",
"tastytrade"
],
"notes": "Table A verdict: MARGINAL - recorded here as feasible_at_100usd=false because the schema boolean admits only an unqualified yes. Minimum position band $5.00-$25.00 (5%-25% of stake); friction band 2.5%-15.0% of stake, renormalized to a stake denominator. The only vehicle offering meaningful implicit leverage a $100 cash account can actually access. Gemini's own JSON recorded feasible=true against its own table verdict of MARGINAL; that defect is not reproduced here."
},
{
"name": "Listed options - micro-options (1-share deliverable)",
"asset_class": "Listed equity options",
"min_viable_position_usd": 0.05,
"min_viable_position_pct_of_stake": 0.05,
"round_trip_friction_pct": null,
"feasible_at_100usd": false,
"regulatory_gates": [
"FINRA Rule 4210(b)(4) $2,000 margin minimum forces a cash account",
"FINRA Rule 2360 options approval - Level 1-2 only (long calls/puts)"
],
"venues": [],
"notes": "Table A verdict: 'Unverified - single source'. MiniMax alone reports a listed micro-option with a 1-share-equivalent deliverable launched 2022-2024, self-tagged [T6] with 'primary citation withheld pending verification'; neither Qwen nor Gemini mentions the instrument. Round-trip friction is Unquantified in every source, hence null. No feasibility verdict in the merged report rests on this row."
},
{
"name": "Listed options - vertical debit spread",
"asset_class": "Listed equity/index options",
"min_viable_position_usd": 5.0,
"min_viable_position_pct_of_stake": 5.0,
"round_trip_friction_pct": 5.0,
"feasible_at_100usd": false,
"regulatory_gates": [
"FINRA Rule 4210(b)(4) $2,000 margin minimum forces a cash account",
"FINRA Rule 2360 options approval Level 3-4 - broker policy generally requires a margin account with $2,000 minimum equity, plus ~30-60 day account seasoning"
],
"venues": [
"Interactive Brokers",
"Charles Schwab",
"tastytrade"
],
"notes": "Table A verdict: GATED. Minimum $5.00-$20.00 net debit; friction band 5.0%-20.0% of stake (two legs of commission plus two legs of spread crossing). Merged finding: structurally cash-securable in principle, practically gated by an approval tier a $100 account is unlikely to clear, compounded by ~30-day Level 3 seasoning against a 90-day clock."
},
{
"name": "Listed options - cash-secured put, $5 strike",
"asset_class": "Listed equity options",
"min_viable_position_usd": 500.0,
"min_viable_position_pct_of_stake": 500.0,
"round_trip_friction_pct": null,
"feasible_at_100usd": false,
"regulatory_gates": [
"FINRA Rule 4210(b)(4) $2,000 margin minimum forces a cash account",
"FINRA Rule 2360 options approval - Level 1-2 only (long calls/puts)"
],
"venues": [],
"notes": "Table A verdict: INFEASIBLE. Requires strike x 100 in cash. At a $1 strike the requirement is exactly the entire stake with maximum gain capped at the premium collected. Friction not stated (n/a) in any source, hence null."
},
{
"name": "Kalshi event contract - P ~ 0.50",
"asset_class": "CFTC-regulated binary event contracts",
"min_viable_position_usd": 0.01,
"min_viable_position_pct_of_stake": 0.01,
"round_trip_friction_pct": 8.0,
"feasible_at_100usd": true,
"regulatory_gates": [
"CFTC designation as a contract market under CEA sec. 5; CEA sec. 5c(c)(5)(C) and 17 CFR 40.11 event-contract prohibition and review",
"Kalshi taker fee formula ceil(0.07 x N x P x (1-P)) per side - exchange rule",
"Massachusetts state gate on non-sports CFTC event contracts (Commonwealth v. KalshiEX LLC, MA Super. Ct. prelim. inj. Jan 2026; MA SJC review pending) - unsettled"
],
"venues": [
"KalshiEX LLC"
],
"notes": "Table A verdict: 'Yes technically; MA contested'. Friction band 8.0%-11.0% of stake. $0.01 is the technical minimum; a doubling attempt requires ~the entire stake. Derived merged result: on a fixed stake the fee collapses to 0.07 x S x (1-P) per side, i.e. 7 x (1-P) dollars per side per $100 - monotonically DECREASING in P, so longshots are the expensive regime and favorites the cheap one. This is the regime where single-shot doubling is mechanically available (P <= 0.50) and where fees consume ~70% of the expected gain on a 5-point edge. Held-to-settlement (no exit trade) is a ~2x uncertainty on every Kalshi figure [T6]."
},
{
"name": "Kalshi event contract - P ~ 0.90 (favorite)",
"asset_class": "CFTC-regulated binary event contracts",
"min_viable_position_usd": 0.01,
"min_viable_position_pct_of_stake": 0.01,
"round_trip_friction_pct": 2.4,
"feasible_at_100usd": true,
"regulatory_gates": [
"CFTC designation as a contract market under CEA sec. 5; CEA sec. 5c(c)(5)(C) and 17 CFR 40.11",
"Kalshi taker fee formula ceil(0.07 x N x P x (1-P)) per side",
"Massachusetts state gate on non-sports CFTC event contracts (Commonwealth v. KalshiEX LLC, MA Super. Ct. prelim. inj. Jan 2026; MA SJC review pending) - unsettled"
],
"venues": [
"KalshiEX LLC"
],
"notes": "Table A verdict: 'Yes technically; MA contested'. Friction band 2.4%-5.4% of stake. Cheap in fee terms but structurally incapable of doubling a stake in one shot: $100 at P=0.90 buys 111 contracts settling at $111. MiniMax's stated verdict that event contracts are feasible for high-confidence (>50%) outcomes to achieve a 100% gross gain is arithmetically impossible and was excluded from the merge."
},
{
"name": "Kalshi event contract - P ~ 0.05 (longshot)",
"asset_class": "CFTC-regulated binary event contracts",
"min_viable_position_usd": 0.01,
"min_viable_position_pct_of_stake": 0.01,
"round_trip_friction_pct": 15.3,
"feasible_at_100usd": false,
"regulatory_gates": [
"CFTC designation as a contract market under CEA sec. 5; CEA sec. 5c(c)(5)(C) and 17 CFR 40.11",
"Kalshi taker fee formula ceil(0.07 x N x P x (1-P)) per side",
"Massachusetts state gate on non-sports CFTC event contracts (Commonwealth v. KalshiEX LLC, MA Super. Ct. prelim. inj. Jan 2026; MA SJC review pending) - unsettled"
],
"venues": [
"KalshiEX LLC"
],
"notes": "Table A verdict: 'No - fee-dominant'. Friction band 15.3%-19.3% of stake. Ceiling rounding alone costs 10% on a single $0.10 contract. The favorite-longshot bias additionally operates against the buyer at this end of the curve."
},
{
"name": "ForecastEx (via IBKR Prediction Markets)",
"asset_class": "CFTC-regulated binary event contracts",
"min_viable_position_usd": 1.0,
"min_viable_position_pct_of_stake": 1.0,
"round_trip_friction_pct": 2.0,
"feasible_at_100usd": true,
"regulatory_gates": [
"CFTC DCM registration - UNVERIFIED; MiniMax concedes it did not retrieve the designation order and Gemini names ForecastEx as lawful without citing its designation at all",
"IBKR account gating; requires a locally running TWS/Gateway process",
"Massachusetts state gate on non-sports CFTC event contracts (Commonwealth v. KalshiEX LLC, MA Super. Ct. prelim. inj. Jan 2026; MA SJC review pending) - unsettled"
],
"venues": [
"ForecastEx LLC via Interactive Brokers"
],
"notes": "Table A verdict: 'Yes technically; MA contested'. Friction band 2.0%-6.0% of stake. Fixed $0.01/contract/side fee makes low-priced contracts disproportionately expensive: ~1.1% per side on a $0.90 contract and ~10% per side on a $0.10 contract. No free historical L2/L3 order-book depth exists for this venue."
},
{
"name": "IBKR CME event contracts",
"asset_class": "Event futures",
"min_viable_position_usd": 0.1,
"min_viable_position_pct_of_stake": 0.1,
"round_trip_friction_pct": 8.0,
"feasible_at_100usd": false,
"regulatory_gates": [
"CFTC DCM (CME Group)",
"IBKR account gating"
],
"venues": [
"Interactive Brokers (CME Group event contracts)"
],
"notes": "Table A verdict: 'No - fixed fee dominant'. Friction band 8.0%-45.0% of stake. Single-source (Gemini): a $0.40 round-trip fixed fee consumes 40% of a $1.00 contract notional."
},
{
"name": "Polymarket (USDC on Polygon)",
"asset_class": "Decentralized prediction market",
"min_viable_position_usd": 0.01,
"min_viable_position_pct_of_stake": 0.01,
"round_trip_friction_pct": 4.0,
"feasible_at_100usd": false,
"regulatory_gates": [
"CFTC order of Jan 3, 2022 ($1.4M, failure to register as a SEF)",
"CFTC Amended Order of Designation reported Nov 2025 following the QCEX acquisition - [T6], Wikipedia-sourced only, no release number cited by any source",
"Massachusetts state gate on non-sports CFTC event contracts (Commonwealth v. KalshiEX LLC, MA Super. Ct. prelim. inj. Jan 2026; MA SJC review pending) - unsettled"
],
"venues": [
"Polymarket"
],
"notes": "Table A verdict: 'No for a MA resident'. Friction band 4.0%-10.0% of stake, dominated by a $1-$5 fiat-to-USDC on-ramp and bridge cost - 1%-5% of a $100 stake and plausibly the largest single line item in Table A. All three reports converge that a Massachusetts resident should not assume lawful access. US retail accessibility itself is genuinely contested: MiniMax reports unblocking on 2025-12-02 [T6], Gemini reports geo-blocking under the 2022 consent order, Qwen reports halt-then-resumption via acquisition of a licensed entity."
},
{
"name": "Spot crypto - advanced / pro order interface",
"asset_class": "Digital assets",
"min_viable_position_usd": 1.0,
"min_viable_position_pct_of_stake": 1.0,
"round_trip_friction_pct": 0.1,
"feasible_at_100usd": true,
"regulatory_gates": [
"FinCEN MSB registration of the venue",
"M.G.L. c. 169 Massachusetts money-transmission regime [T6] - Massachusetts has not adopted a distinct crypto licensing regime and does not prohibit crypto trading"
],
"venues": [
"Coinbase Advanced",
"Kraken Pro"
],
"notes": "Table A verdict: Yes. Friction band 0.10%-0.60% of stake (0.05%-0.60% maker/taker by volume tier plus 0.10%-0.50% spread). Two-source majority. The retail simple-trade interface at the same venues costs roughly 4x more for identical exposure - a pure interface-selection cost with no offsetting benefit."
},
{
"name": "Spot crypto - retail 'simple trade' interface",
"asset_class": "Digital assets",
"min_viable_position_usd": 1.0,
"min_viable_position_pct_of_stake": 1.0,
"round_trip_friction_pct": 0.8,
"feasible_at_100usd": true,
"regulatory_gates": [
"FinCEN MSB registration of the venue",
"M.G.L. c. 169 Massachusetts money-transmission regime [T6]"
],
"venues": [
"Coinbase (simple trade)",
"Kraken (instant buy)"
],
"notes": "Table A verdict: 'Yes, ~4x the pro-interface cost'. Friction band 0.80%-2.50% of stake, cost embedded in the spread rather than charged as a fee. Carried as a distinct row rather than pruned because the two regimes differ by roughly 4x and the distinction is decision-relevant."
},
{
"name": "Crypto nano / micro futures",
"asset_class": "Crypto futures",
"min_viable_position_usd": 20.0,
"min_viable_position_pct_of_stake": 20.0,
"round_trip_friction_pct": 3.5,
"feasible_at_100usd": false,
"regulatory_gates": [
"CFTC DCM",
"Margin equity collateral $20-$50 minimum"
],
"venues": [
"Coinbase Derivatives"
],
"notes": "Table A verdict: No. Minimum position $20.00-$50.00 (20%-50% of stake); friction band 3.5%-12.0%. Gemini and MiniMax converge on infeasible; Qwen's claim of feasibility with a 'fraction of a cent' minimum position is incoherent for a margined futures contract and was pruned. High liquidation risk on a $100 balance."
},
{
"name": "CME Bitcoin futures (standard, 5 BTC)",
"asset_class": "Crypto futures",
"min_viable_position_usd": 200000.0,
"min_viable_position_pct_of_stake": 200000.0,
"round_trip_friction_pct": null,
"feasible_at_100usd": false,
"regulatory_gates": [
"CFTC DCM (CME Group)",
"$2,000 minimum margin account plus $200,000-$260,000 initial margin"
],
"venues": [
"CME Group via Interactive Brokers"
],
"notes": "Table A verdict: 'Decisively infeasible'. ~$575,000 notional (5 BTC); minimum viable position recorded as the low end of the $200,000-$260,000 initial-margin band. Commission is irrelevant at this scale so no friction percentage is stated by any source, hence null. MiniMax gives the notional three incompatible values across its own document (~$5,000 headline; $575,000 body and table; $200,000 elsewhere); the infeasibility verdict is robust to the entire range - even at a $2,000 initial margin the vehicle is 20x the stake."
},
{
"name": "KalshiEX BTCPERP (perpetual)",
"asset_class": "Crypto perpetual futures on a CFTC DCM",
"min_viable_position_usd": 1.0,
"min_viable_position_pct_of_stake": 1.0,
"round_trip_friction_pct": 0.25,
"feasible_at_100usd": false,
"regulatory_gates": [
"CFTC approval reported at release 9240-26 (May 29, 2026) - [T6], the entire 92xx-26 release series is unverified against primary source",
"Massachusetts state gate on non-sports CFTC event contracts (Commonwealth v. KalshiEX LLC, MA Super. Ct. prelim. inj. Jan 2026; MA SJC review pending) - unsettled"
],
"venues": [
"KalshiEX LLC"
],
"notes": "Table A verdict: 'Marginal - leverage-binding, single source'. Friction band 0.25%-15.0% on margin deployed. Single-source [T6]; MiniMax concedes the contract specification was not directly retrieved. Fee is charged on full position notional rather than posted margin, so at 5x-50x leverage the effective drag on deployed capital is 5x-50x the notional rate."
}
],
"strategies": [
{
"name": "Market excess return (Mkt-RF)",
"vehicle": "Fractional equity / ETF",
"evidence_tier": "T1",
"citations": [
{
"title": "Capital Asset Prices (Sharpe 1964)",
"doi_or_url": "10.1111/j.1540-6261.1964.tb02865.x",
"year": 1964
},
{
"title": "Common Risk Factors in the Returns on Stocks and Bonds (Fama & French 1993)",
"doi_or_url": "10.1016/0304-405X(93)90023-5",
"year": 1993
}
],
"documented_effect_size": "~6-8% annualized real",
"post_publication_decay": "Low; no literature argues it has disappeared",
"min_capital_usd": 1.0,
"p_reach_target": null,
"p_ruin": null,
"median_days_to_target": null,
"feasible": false,
"parameters": {
"friction_breakeven": "<0.1% (spread only)",
"note": "median_days_to_target and its IQR are null for every strategy: Section 12 resolved that neither MiniMax nor Gemini derives a first-passage TIME distribution anywhere - both derive only terminal-return distributions - so filling these cells would be fabrication.",
"table_c_feasibility_verdict": "No - horizon; ~1.7% expected over 90d"
}
},
{
"name": "Size (SMB)",
"vehicle": "Fractional equity / ETF",
"evidence_tier": "T2",
"citations": [
{
"title": "Common Risk Factors in the Returns on Stocks and Bonds (Fama & French 1993)",
"doi_or_url": "10.1016/0304-405X(93)90023-5",
"year": 1993
}
],
"documented_effect_size": "~2% annualized",
"post_publication_decay": "Modest; contested post-1980",
"min_capital_usd": 1000.0,
"p_reach_target": null,
"p_ruin": null,
"median_days_to_target": null,
"feasible": false,
"parameters": {
"friction_breakeven": "n/a at this scale",
"table_c_feasibility_verdict": "No - horizon-infeasible"
}
},
{
"name": "Value (HML)",
"vehicle": "Fractional equity / ETF",
"evidence_tier": "T2 contested",
"citations": [
{
"title": "Common Risk Factors in the Returns on Stocks and Bonds (Fama & French 1993)",
"doi_or_url": "10.1016/0304-405X(93)90023-5",
"year": 1993
},
{
"title": "A Five-Factor Asset Pricing Model (Fama & French 2015)",
"doi_or_url": "10.1016/j.jfineco.2014.10.010",
"year": 2015
}
],
"documented_effect_size": "~0-3% annualized",
"post_publication_decay": "Severe 2017-2020; partial post-2021 rebound in higher-inflation regimes; rejected as an independent factor by the q-model once investment and profitability are included",
"min_capital_usd": 1000.0,
"p_reach_target": null,
"p_ruin": null,
"median_days_to_target": null,
"feasible": false,
"parameters": {
"friction_breakeven": "~0.1% per rebalance",
"table_c_feasibility_verdict": "No - sub-1% over 90d"
}
},
{
"name": "Momentum, cross-sectional 3-12m",
"vehicle": "Fractional equity / ETF",
"evidence_tier": "T1 effect / T2 magnitude",
"citations": [
{
"title": "Returns to Buying Winners and Selling Losers (Jegadeesh & Titman 1993)",
"doi_or_url": "10.1111/j.1540-6261.1993.tb04702.x",
"year": 1993
},
{
"title": "Value and Momentum Everywhere (Asness, Moskowitz & Pedersen 2013)",
"doi_or_url": "10.1111/jofi.12021",
"year": 2013
}
],
"documented_effect_size": "4-8% annualized post-decay (~0.5%/mo)",
"post_publication_decay": "~50-58% (McLean & Pontiff)",
"min_capital_usd": 10.0,
"p_reach_target": 0.01,
"p_ruin": null,
"median_days_to_target": null,
"feasible": false,
"parameters": {
"friction_breakeven": "0.5-1% per rebalance budgeted; 4-8% actual at retail",
"total_friction_drag_pct_of_stake": "0.25-10%",
"after_tax_ev_delta_usd": "-8 to -21",
"crash_risk": "Severely left-skewed return distribution; 2008-2009 momentum crash [T1]",
"citation_correction": "Gemini attributed cross-sectional momentum to Harvey, Liu & Zhu (2016), a multiple-testing critique, not a momentum result; misattribution dropped and replaced.",
"table_c_feasibility_verdict": "No - friction exceeds budget 8-16x; 90d too short",
"p_ruin_band": "0.20-0.35",
"p_ruin_definition": "Section 12 reports P(ruin) only as a band, never as a row-level central, so p_ruin is null rather than a midpoint. 'Ruin' throughout means terminal wealth <= USD 25 (an experiment-killing loss), NOT literal total loss.",
"table_g_rank": 6
}
},
{
"name": "Time-series momentum",
"vehicle": "Fractional equity / ETF",
"evidence_tier": "T1",
"citations": [
{
"title": "Time Series Momentum (Moskowitz, Ooi & Pedersen 2012)",
"doi_or_url": "10.1016/j.jfineco.2011.11.003",
"year": 2012
}
],
"documented_effect_size": "Comparable to cross-sectional; long-only",
"post_publication_decay": "Not separately quantified in the merged corpus",
"min_capital_usd": 10.0,
"p_reach_target": null,
"p_ruin": null,
"median_days_to_target": null,
"feasible": false,
"parameters": {
"friction_breakeven": "0.5% per rebalance",
"table_c_feasibility_verdict": "No - same horizon and friction failure"
}
},
{
"name": "Short-term reversal (1-week)",
"vehicle": "Fractional equity - liquid single names",
"evidence_tier": "T1 effect / T2 magnitude",
"citations": [
{
"title": "Evidence of Predictable Behavior of Security Returns (Jegadeesh 1990)",
"doi_or_url": "10.1111/j.1540-6261.1990.tb05110.x",
"year": 1990
},
{
"title": "Fads, Martingales, and Market Efficiency (Lehmann 1990)",
"doi_or_url": "10.2307/2330889",
"year": 1990
}
],
"documented_effect_size": "0.5-1.0% per week gross",
"post_publication_decay": "Substantial; compressed as electronic market-making absorbed the liquidity-provision return",
"min_capital_usd": 50.0,
"p_reach_target": 0.01,
"p_ruin": null,
"median_days_to_target": null,
"feasible": false,
"parameters": {
"friction_breakeven": "$500 equivalent; 80%+ of gross consumed by fees",
"total_friction_drag_pct_of_stake": "5-12%",
"after_tax_ev_delta_usd": "-8 to -18",
"author_derivation": "Gross 0.5-1.0%/wk over 13 weeks = 6.5-13%, against friction of 4-12% over the same window; expected net bounded at -5.5% to +9%, full cross-range -11.5% to +9% [T6]",
"table_c_feasibility_verdict": "No - net over 90d straddles zero (-5.5% to +9%)",
"p_ruin_band": "0.25-0.40",
"p_ruin_definition": "Section 12 reports P(ruin) only as a band, never as a row-level central, so p_ruin is null rather than a midpoint. 'Ruin' throughout means terminal wealth <= USD 25 (an experiment-killing loss), NOT literal total loss.",
"table_g_rank": 7
}
},
{
"name": "Profitability (RMW)",
"vehicle": "Fractional equity / ETF",
"evidence_tier": "T1",
"citations": [
{
"title": "The Other Side of Value: The Gross Profitability Premium (Novy-Marx 2013)",
"doi_or_url": "10.1016/j.jfineco.2013.01.003",
"year": 2013
}
],
"documented_effect_size": "~0.5%/month (magnitude is [T6], single-sourced)",
"post_publication_decay": "Modest",
"min_capital_usd": 300.0,
"p_reach_target": null,
"p_ruin": null,
"median_days_to_target": null,
"feasible": false,
"parameters": {
"friction_breakeven": "~0.3% per rebalance",
"table_c_feasibility_verdict": "No - ~1.5% over 90d"
}
},
{
"name": "Investment (CMA)",
"vehicle": "Fractional equity / ETF",
"evidence_tier": "T1",
"citations": [
{
"title": "Asset Growth and the Cross-Section of Stock Returns (Cooper, Gulen & Schill 2008)",
"doi_or_url": "10.1111/j.1540-6261.2008.01369.x",
"year": 2008
},
{
"title": "Capital Investments and Stock Returns (Titman, Wei & Xie 2004)",
"doi_or_url": "10.1017/S0022109000003125",
"year": 2004
}
],
"documented_effect_size": "~0.3%/month (magnitude is [T6])",
"post_publication_decay": "Modest",
"min_capital_usd": 300.0,
"p_reach_target": null,
"p_ruin": null,
"median_days_to_target": null,
"feasible": false,
"parameters": {
"friction_breakeven": "~0.3% per rebalance",
"table_c_feasibility_verdict": "No"
}
},
{
"name": "Quality (QMJ)",
"vehicle": "Fractional equity / ETF",
"evidence_tier": "T1 effect / T2 independence",
"citations": [
{
"title": "Quality Minus Junk (Asness, Frazzini & Pedersen 2019)",
"doi_or_url": "10.1007/s11142-018-9470-2",
"year": 2019
}
],
"documented_effect_size": "~0.4%/month (magnitude is [T6])",
"post_publication_decay": "Modest",
"min_capital_usd": 300.0,
"p_reach_target": null,
"p_ruin": null,
"median_days_to_target": null,
"feasible": false,
"parameters": {
"friction_breakeven": "~0.3% per rebalance",
"citation_correction": "MiniMax cited 'Asness, Frazzini & Israel (2019), AQR Working Paper' at [T4]; wrong third author, wrong venue, and wrong tier for a peer-reviewed paper.",
"table_c_feasibility_verdict": "No"
}
},
{
"name": "Betting Against Beta (BAB)",
"vehicle": "Fractional equity - requires shorting",
"evidence_tier": "T1; subsumption claim T6",
"citations": [
{
"title": "Betting Against Beta (Frazzini & Pedersen 2014)",
"doi_or_url": "10.1016/j.jfineco.2013.10.005",
"year": 2014
}
],
"documented_effect_size": "~0.5%/month ([T6]); premium peaks in financial stress",
"post_publication_decay": "Contested - subsumption by standard risk factors claimed but unverified",
"min_capital_usd": 1000.0,
"p_reach_target": null,
"p_ruin": null,
"median_days_to_target": null,
"feasible": false,
"parameters": {
"friction_breakeven": "0.5% per rebalance",
"hard_gate": "Requires shorting; margin account unavailable at $100",
"table_c_feasibility_verdict": "No - requires shorting; margin account unavailable"
}
},
{
"name": "Idiosyncratic / low volatility",
"vehicle": "Fractional equity / ETF",
"evidence_tier": "T2",
"citations": [],
"documented_effect_size": "~0.4%/month ([T6])",
"post_publication_decay": "Modest; material trading-cost deduction",
"min_capital_usd": 300.0,
"p_reach_target": null,
"p_ruin": null,
"median_days_to_target": null,
"feasible": false,
"parameters": {
"friction_breakeven": "~0.5% per rebalance",
"citation_status": "NO key citation is asserted. MiniMax's sole citation could not be located with the authors given and one named author is a market-structure rather than an idiosyncratic-volatility researcher, suggesting confabulated authorship; neither Gemini nor Qwen covers the factor. The anomaly is genuinely well-established in the wider literature but no citation in this corpus is trustworthy enough to attach, and none was invented.",
"table_c_feasibility_verdict": "No"
}
},
{
"name": "Carry (FX / bond / commodity)",
"vehicle": "Futures - institutional",
"evidence_tier": "T1",
"citations": [
{
"title": "Carry (Koijen, Moskowitz, Pedersen & Vrugt 2018)",
"doi_or_url": "10.1016/j.jfineco.2017.11.002",
"year": 2018
}
],
"documented_effect_size": "4-8% annualized",
"post_publication_decay": "Modest; survives across asset classes",
"min_capital_usd": 10000.0,
"p_reach_target": null,
"p_ruin": null,
"median_days_to_target": null,
"feasible": false,
"parameters": {
"friction_breakeven": "n/a at this scale",
"table_c_feasibility_verdict": "No - institutional infrastructure required"
}
},
{
"name": "Post-earnings-announcement drift (PEAD)",
"vehicle": "Fractional equity - commission-free",
"evidence_tier": "T1 effect / T2 magnitude",
"citations": [
{
"title": "An Empirical Evaluation of Accounting Income Numbers (Ball & Brown 1968)",
"doi_or_url": "10.2307/2490232",
"year": 1968
},
{
"title": "Post-Earnings-Announcement Drift (Bernard & Thomas 1989), JAR 27 Supplement",
"doi_or_url": "10.2307/2491256",
"year": 1989
},
{
"title": "Evidence that Stock Prices Do Not Fully Reflect the Implications of Current Earnings (Bernard & Thomas 1990)",
"doi_or_url": "10.1016/0165-4101(90)90008-R",
"year": 1990
}
],
"documented_effect_size": "+2% to +5% on the long leg over a 30-day hold, top-decile SUE",
"post_publication_decay": "~35-50%; the decay citation in one source does not support the claim (it studies profitability and book-to-market, not PEAD), so the decay is real, directionally large, and imprecisely measured",
"min_capital_usd": 5.0,
"p_reach_target": 0.01,
"p_ruin": null,
"median_days_to_target": null,
"feasible": false,
"parameters": {
"sue_threshold": "> 2.0 standard deviations",
"holding_period_days": 30,
"max_positions": 5,
"friction_breakeven": "~0.3% per round trip; needs sub-5bp spreads",
"total_friction_drag_pct_of_stake": "1-12% (driver is trade count: 0.05% is one round trip, 4-12% assumes 20-40 trades)",
"after_tax_ev_delta_usd": "-12 to -25",
"author_derivation": "At +2% to +5% per 30-day event with a maximum of three sequential holds in 90 days the compounded range is +6% to +16% before friction [T6]",
"mechanism_note": "A follow-up working paper suggests modern PEAD is concentrated in stocks with no sell-side analyst following, making the residual a limited-attention premium [T3], single-sourced - which would place it in exactly the low-liquidity names where retail friction is worst.",
"table_c_feasibility_verdict": "Marginal on execution, No on target - <= ~16% compounded over 3 events",
"p_ruin_band": "0.15-0.30",
"p_ruin_definition": "Section 12 reports P(ruin) only as a band, never as a row-level central, so p_ruin is null rather than a midpoint. 'Ruin' throughout means terminal wealth <= USD 25 (an experiment-killing loss), NOT literal total loss.",
"table_g_rank": 5
}
},
{
"name": "Long-term reversal (3-5y)",
"vehicle": "Fractional equity / ETF",
"evidence_tier": "T1",
"citations": [
{
"title": "Does the Stock Market Overreact? (DeBondt & Thaler 1985)",
"doi_or_url": "10.1111/j.1540-6261.1985.tb05004.x",
"year": 1985
}
],
"documented_effect_size": "~5% annualized",
"post_publication_decay": "Significant",
"min_capital_usd": 1000.0,
"p_reach_target": null,
"p_ruin": null,
"median_days_to_target": null,
"feasible": false,
"parameters": {
"friction_breakeven": "0.5% per rebalance",
"table_c_feasibility_verdict": "No - horizon exceeds mandate by 12-20x"
}
},
{
"name": "Accruals",
"vehicle": "Fractional equity / ETF",
"evidence_tier": "T2",
"citations": [
{
"title": "Do Stock Prices Fully Reflect Information in Accruals and Cash Flows About Future Earnings? (Sloan 1996), Accounting Review 71(3), 289-315",
"doi_or_url": null,
"year": 1996
}
],
"documented_effect_size": "~2-4% annualized, materially reduced",
"post_publication_decay": "Substantial; partly subsumed by profitability",
"min_capital_usd": 1000.0,
"p_reach_target": null,
"p_ruin": null,
"median_days_to_target": null,
"feasible": false,
"parameters": {
"friction_breakeven": "~0.3% per rebalance",
"doi_status": "DOI disputed across sources; none asserted",
"table_c_feasibility_verdict": "No"
}
},
{
"name": "Net stock issuance",
"vehicle": "Fractional equity / ETF",
"evidence_tier": "T2",
"citations": [
{
"title": "The New Issues Puzzle (Loughran & Ritter 1995), Journal of Finance 50(1)",
"doi_or_url": null,
"year": 1995
}
],
"documented_effect_size": "~2-4% annualized",
"post_publication_decay": "Substantial",
"min_capital_usd": 1000.0,
"p_reach_target": null,
"p_ruin": null,
"median_days_to_target": null,
"feasible": false,
"parameters": {
"friction_breakeven": "~0.3% per rebalance",
"doi_status": "DOI internally inconsistent in the source; not asserted",
"table_c_feasibility_verdict": "No"
}
},
{
"name": "Volatility risk premium - short premium",
"vehicle": "Listed options - short premium",
"evidence_tier": "T1 effect / T2 magnitude",
"citations": [
{
"title": "Variance Risk Premiums (Carr & Wu 2009)",
"doi_or_url": "10.1093/rfs/hhn038",
"year": 2009
},
{
"title": "Expected Option Returns (Coval & Shumway 2001)",
"doi_or_url": "10.1111/0022-1082.00352",
"year": 2001
}
],
"documented_effect_size": "+1% to +2% per month ([T2], single-sourced); crash risk -30% to -90% of portfolio value in a single session (5 Feb 2018, Mar 2020) [T4]",
"post_publication_decay": "Narrowed post-2014; still positive",
"min_capital_usd": 2000.0,
"p_reach_target": null,
"p_ruin": null,
"median_days_to_target": null,
"feasible": false,
"parameters": {
"friction_breakeven": "0.5-2% per round trip on liquid SPX/SPY",
"hard_gate": "$2,000 cash-secured-put margin floor; $10,000+ for naked/spread writing at Level 3/4 [T5]. 20x-100x the stake.",
"struck_figures": "MiniMax's ~0.10%/day (prose) and ~0.5%/day (table) magnitudes and its 0.50-0.75 short-straddle Sharpe attributed to Coval & Shumway are struck - that paper reports large NEGATIVE straddle returns of roughly -3%/week.",
"framing": "The VRP is compensation for bearing crash risk, not a free lunch [T1]. Under a first-passage objective a strategy whose left tail can remove 90% of capital in one session is worse than its Sharpe suggests - ruin is absorbing.",
"table_c_feasibility_verdict": "No - capital and approval-tier gated 20x above stake"
}
},
{
"name": "Long premium / long volatility (single-leg directional)",
"vehicle": "Listed options - long premium",
"evidence_tier": "T1 sign",
"citations": [
{
"title": "Expected Option Returns (Coval & Shumway 2001)",
"doi_or_url": "10.1111/0022-1082.00352",
"year": 2001
},
{
"title": "Retail Trading in Options and the Rise of the Big Three Wholesalers (Bryzgalova, Pavlova & Sikorskaya 2023), Journal of Finance 78(6)",
"doi_or_url": null,
"year": 2023
}
],
"documented_effect_size": "Negative EV; the -5% to -10% annualized magnitude is [T6] and uncorroborated; Gemini's -15% to -30% per trade for the 0DTE variant is [T6], sign corroborated, magnitude not",
"post_publication_decay": "n/a",
"min_capital_usd": 5.0,
"p_reach_target": 0.02,
"p_ruin": null,
"median_days_to_target": null,
"feasible": false,
"parameters": {
"friction_breakeven": "~0.5% per round trip plus a 10-30% spread cross",
"total_friction_drag_pct_of_stake": "2.5-12%",
"after_tax_ev_delta_usd": "-20 to -58",
"shared_table_g_row": "Table G row 4 covers 'Long premium directional options, single-leg (incl. full-stake bold play and 0DTE variants)' and its figures are attached to both this row and the 0DTE row; they are one merged row, not two independent estimates.",
"band_note": "P(reach) band 0.01-0.08 is genuinely unresolved - MiniMax puts 0DTE below 0.01, Qwen puts deep-OTM buying below 0.05, Gemini puts long premium at 0.08, and none models it.",
"structural_tension": "Under Dubins-Savage, bold play maximizes P(reach) in a subfair game and a single OTM call doubles on a far smaller underlying move than the underlying's own doubling requires [T1]. That is precisely why this row also carries the table's worst P(ruin) and worst after-tax EV. Under a fixed-multiple fixed-deadline objective the strategy with positive EV cannot reach the target and the strategy that can reach the target has negative EV.",
"table_c_feasibility_verdict": "No - structurally on the wrong side of the VRP",
"p_ruin_band": "0.55-0.92",
"p_ruin_definition": "Section 12 reports P(ruin) only as a band, never as a row-level central, so p_ruin is null rather than a midpoint. 'Ruin' throughout means terminal wealth <= USD 25 (an experiment-killing loss), NOT literal total loss.",
"table_g_rank": 4
}
},
{
"name": "0DTE long directional (retail)",
"vehicle": "Listed options - 0DTE",
"evidence_tier": "T3/T4; magnitude T6",
"citations": [
{
"title": "Retail Trading in Options and the Rise of the Big Three Wholesalers (Bryzgalova, Pavlova & Sikorskaya 2023), Journal of Finance 78(6)",
"doi_or_url": null,
"year": 2023
},
{
"title": "Cboe exchange volume data [T4]; SSRN working papers 2023-2025 [T3]",
"doi_or_url": null,
"year": 0
}
],
"documented_effect_size": "Negative EV; -15% to -30% per trade [T6]; retail loses 65-80% of premium over 12-month windows [T4], unnamed industry source",
"post_publication_decay": "n/a - the market is post-2022",
"min_capital_usd": 10.0,
"p_reach_target": 0.02,
"p_ruin": null,
"median_days_to_target": null,
"feasible": false,
"parameters": {
"bid_ask_drag": "10-30% of premium",
"theta_drag": "extreme",
"shared_table_g_row": "Shares Table G row 4 with the long-premium row above; the probabilities are one merged estimate covering both.",
"struck_figures": "MiniMax's 'retail loses 0.5-1.5% of premium per trade' is struck as arithmetically irreconcilable with a sub-10% win rate on capped-loss instruments. The '<10% of trades profitable' figure is [T6] - unverifiable rather than fabricated (poorly indexed venue).",
"gamma_note": "0DTE gamma exposure is a documented intraday volatility-SUPPRESSION mechanism [T3]; the 0DTE VRP is reportedly LARGER than the standard SPX VRP [T3], which makes retail's position worse rather than better since retail is the buyer.",
"table_c_feasibility_verdict": "No - median outcome is total loss of premium",
"p_ruin_band": "0.55-0.92",
"p_ruin_definition": "Section 12 reports P(ruin) only as a band, never as a row-level central, so p_ruin is null rather than a midpoint. 'Ruin' throughout means terminal wealth <= USD 25 (an experiment-killing loss), NOT literal total loss.",
"table_g_rank": 4
}
},
{
"name": "Retail options trading, general",
"vehicle": "Listed options",
"evidence_tier": "T1 direction / T2 magnitude",
"citations": [
{
"title": "Retail Trading in Options and the Rise of the Big Three Wholesalers (Bryzgalova, Pavlova & Sikorskaya 2023), Journal of Finance 78(6)",
"doi_or_url": null,
"year": 2023
},
{
"title": "Attention-Induced Trading and Returns: Evidence from Robinhood Users (Barber, Huang, Odean & Schwarz 2022), Journal of Finance 77(6), 3141-3190",
"doi_or_url": null,
"year": 2022
}
],
"documented_effect_size": "Negative EV in the population average; '<10% of trades profitable' is [T6], single-sourced and unverified",
"post_publication_decay": "n/a",
"min_capital_usd": 100.0,
"p_reach_target": null,
"p_ruin": null,
"median_days_to_target": null,
"feasible": false,
"parameters": {
"friction_breakeven": "Variable; spread plus theta",
"provenance_note": "Barber, Huang, Odean & Schwarz (2022) was surfaced by a verification pass as a canonical omission, not asserted by any of the three report bodies. No DOI is asserted because none appears anywhere in the corpus and none was invented.",
"table_c_feasibility_verdict": "No - structurally lossy"
}
},
{
"name": "Merger arbitrage / event-driven",
"vehicle": "Fractional equity",
"evidence_tier": "T1",
"citations": [
{
"title": "Characteristics of Risk and Return in Risk Arbitrage (Mitchell & Pulvino 2001)",
"doi_or_url": "10.1111/0022-1082.00418",
"year": 2001
}
],
"documented_effect_size": "2-6% annualized; 1-3% per low-risk deal over 30-60 days; a broken deal loses -30% to -50% in a day",
"post_publication_decay": "Narrowed as event-driven funds crowded in",
"min_capital_usd": 1000.0,
"p_reach_target": null,
"p_ruin": null,
"median_days_to_target": null,
"feasible": false,
"parameters": {
"friction_breakeven": "~0.5% per round trip; $500 breakeven",
"position_vs_strategy": "$1,000 is the capital needed for the STRATEGY (multiple deals); ~$100 funds a single micro-lot position. The two figures answer different questions.",
"payoff": "At $100 there is no deal diversification - the position is a single binary with an asymmetric payoff of roughly +2% versus -40%. Reaching the target requires being right about a deal BREAKING, which is the short side and is not what the risk-arbitrage literature documents.",
"citation_correction": "MiniMax's sole [T1] merger-arbitrage source could not be located and appears fabricated; its supporting citations included a venture-capital valuation-waterfall paper with no bearing on merger spreads.",
"table_c_feasibility_verdict": "No - single-deal binary; payoff +2% vs -40%"
}
},
{
"name": "Index reconstitution arbitrage",
"vehicle": "Fractional equity",
"evidence_tier": "T1 effect / T2 magnitude",
"citations": [
{
"title": "Price and Volume Effects Associated with Changes in the S&P 500 List (Harris & Gurel 1986)",
"doi_or_url": "10.1111/j.1540-6261.1986.tb04550.x",
"year": 1986
},
{
"title": "Do Demand Curves for Stocks Slope Down? (Shleifer 1986)",
"doi_or_url": "10.1111/j.1540-6261.1986.tb04518.x",
"year": 1986
},
{
"title": "Does Arbitrage Flatten Demand Curves for Stocks? (Wurgler & Zhuravskaya 2002)",
"doi_or_url": "10.1086/341638",
"year": 2002
}
],
"documented_effect_size": "1.5-3.0% per event, decayed from the 1986-era effect",
"post_publication_decay": "Substantial - ETF-driven arbitrage",
"min_capital_usd": 100.0,
"p_reach_target": null,
"p_ruin": null,
"median_days_to_target": null,
"feasible": false,
"parameters": {
"friction_breakeven": "~0.3-1.5% per round trip; net alpha ~0.5-1.5%",
"calendar_gate": "The Russell reconstitution is annual, in June [T5]. A 90-day window either contains a reconstitution event or it does not, and if it does not the strategy has zero trading opportunities.",
"excluded_figures": "MiniMax supplied four mutually incompatible magnitudes for the same effect within one document (0.2-0.3%, 2-4%, 1-2%, 2-4%); disqualified as internally incoherent. Gemini's Madhavan (2003) DOI carries a JPM prefix for an FAJ article and is not asserted.",
"table_c_feasibility_verdict": "No - expected return insufficient; Russell reconstitution is annual (June)"
}
},
{
"name": "Prediction market - favorite buying (favorite-longshot-bias harvest)",
"vehicle": "Kalshi / ForecastEx event contracts",
"evidence_tier": "T1 phenomenon / T6 tradeable magnitude",
"citations": [
{
"title": "Explaining the Favorite-Longshot Bias: Is it Risk-Love or Misperceptions? (Snowberg & Wolfers 2010), JPE 118(4), 723-746",
"doi_or_url": null,
"year": 2010
},
{
"title": "The Economics of Wagering Markets (Sauer 1998), JEL 36(4), 2021-2064",
"doi_or_url": null,
"year": 1998
}
],
"documented_effect_size": "TRANSFERRED (racetrack): longshots overpriced ~25-30%, favorites underpriced ~3-5% [T1]. The claimed DIRECT Kalshi figure of +5% to +12% EV at P>=0.70 is [T6] and was REJECTED as an uncorroborated report harmonization.",
"post_publication_decay": "Persistent in retail-dominated venues [T2]; no decay series exists for regulated event contracts",
"min_capital_usd": 1.0,
"p_reach_target": 0.035,
"p_ruin": null,
"median_days_to_target": null,
"feasible": false,
"parameters": {
"target_contract_odds": "P >= 0.70 (as specified by the source that claimed the effect)",
"order_type": "Maker limit order - 'maker = 0% fee' is [T6], uncited, and load-bearing for the rejected magnitude",
"max_trade_allocation_pct": 20.0,
"kalshi_taker_fee": "ceil(0.07 x P x (1-P) x N)/100 per side; ~10% drag on a $0.10 bet [T5]",
"total_friction_drag_pct_of_stake": "3.6-15%",
"after_tax_ev_delta_usd": "-8 to -18",
"legal_condition": "MA access is contested. Table G row 3 is conditional on the Massachusetts question resolving favorably or on the participant not residing in Massachusetts. The merged headline survives this row collapsing entirely because the ceiling is crypto-driven.",
"author_derivation": "At P=0.70 doubling needs two consecutive full-stake wins (1.43^2 = 2.04), which occurs with probability 0.70^2 = 49% in a FAIR market against ~51% ruin - a coin flip with no edge, which is what a correctly priced contract should deliver. Any excess over 49% must come entirely from the FLB edge, transferred at 3-5%, not the 5-12% claimed. [T6]",
"doi_status": "MiniMax gives 10.1086/655443 and Gemini gives 10.1086/655844 for Snowberg & Wolfers; irreconcilable on available evidence and NEITHER DOI is asserted.",
"table_c_feasibility_verdict": "Legally contingent - deferred. On economics alone: No",
"p_ruin_band": "0.30-0.45",
"p_ruin_definition": "Section 12 reports P(ruin) only as a band, never as a row-level central, so p_ruin is null rather than a midpoint. 'Ruin' throughout means terminal wealth <= USD 25 (an experiment-killing loss), NOT literal total loss.",
"table_g_rank": 3
}
},
{
"name": "Prediction market - informed / asymmetric-information trading",
"vehicle": "Kalshi / ForecastEx event contracts",
"evidence_tier": "T2",
"citations": [
{
"title": "Wolfers & Zitzewitz, Journal of Economic Perspectives - year, volume and DOI disputed across sources; not asserted",
"doi_or_url": null,
"year": 0
}
],
"documented_effect_size": "Positive EV documented for WELL-INFORMED traders; four enabling conditions, two of which exclude a $100 account",
"post_publication_decay": "n/a",
"min_capital_usd": 100.0,
"p_reach_target": 0.035,
"p_ruin": null,
"median_days_to_target": null,
"feasible": false,
"parameters": {
"enabling_conditions": "(1) persistent miscalibration; (2) categories with structural information asymmetry; (3) markets illiquid enough that a single position moves the price; (4) high capital and rapid execution. Conditions (3) and (4) are mutually hostile at $100 - a stake that can move a thin market is a stake that cannot exit it - and (4) explicitly excludes the subject of this report.",
"shared_table_g_row": "Shares Table G row 3 with the favorite-buying row.",
"evidence_absence": "Direct peer-reviewed studies of retail-account profitability on Kalshi, ForecastEx or Polymarket are practically nonexistent as of 2026-08-01.",
"table_c_feasibility_verdict": "No - documented edge accrues to high-capital, fast-execution informed traders",
"p_ruin_band": "0.30-0.45",
"p_ruin_definition": "Section 12 reports P(ruin) only as a band, never as a row-level central, so p_ruin is null rather than a midpoint. 'Ruin' throughout means terminal wealth <= USD 25 (an experiment-killing loss), NOT literal total loss.",
"table_g_rank": 3
}
},
{
"name": "Prediction market - retail profitability, generally",
"vehicle": "Kalshi / ForecastEx / Polymarket",
"evidence_tier": "T1 for the ABSENCE of evidence",
"citations": [],
"documented_effect_size": "No documented effect size exists. Grey-lit signals: 'up to 80% of users are net losers' [T4]; 'top 1% capture 84% of gains' [T4], with a platform mismatch flagged (attributed to Kalshi, sourced to two Polymarket references).",
"post_publication_decay": "n/a",
"min_capital_usd": null,
"p_reach_target": null,
"p_ruin": null,
"median_days_to_target": null,
"feasible": false,
"parameters": {
"min_capital_usd_is_inapplicable_not_unknown": "Table C records this row's minimum-viable-capital cell as 'n/a', not as an unmeasured number. This row is an EVIDENCE-ABSENCE FINDING about the state of the literature, not an executable strategy with a capital floor, so the field does not apply. The null must not be read as missing data - the two sibling prediction-market rows carry 100.0 and 0.01 respectively.",
"citation_status": "No peer-reviewed study exists as of 2026-08-01 quantifying retail Sharpe or hit rates on Kalshi, ForecastEx or Polymarket. MiniMax's Polymarket FLB of ~5-10% and retail informed-trader returns of 1-5% per trade were STRUCK ENTIRELY - both rest solely on 'Penn, C. (2025), An Empirical Study of Prediction Markets, forthcoming International Journal of Forecasting', a paper that does not exist.",
"table_c_feasibility_verdict": "No - any positive claim is unsupported by the peer-reviewed literature"
}
},
{
"name": "Spot crypto held outright (BTC or comparable high-volatility major), weekly rebalance, no leverage",
"vehicle": "Spot crypto - advanced/pro order interface",
"evidence_tier": "T1 realized-return distribution / T5 venue fee schedules / T6 probabilities",
"citations": [],
"documented_effect_size": "Not expressed as a documented anomaly effect size; the row rests on the realized cross-sectional volatility of the asset class rather than on a published premium",
"post_publication_decay": "n/a - not a published anomaly",
"min_capital_usd": 1.0,
"p_reach_target": 0.05,
"p_ruin": null,
"median_days_to_target": null,
"feasible": false,
"parameters": {
"total_friction_drag_pct_of_stake": "0.8-3.0%",
"after_tax_ev_delta_usd": "-12 to -29",
"p_reach_band": "0.03-0.08",
"coherence_warning": "This row's band is the least coherently derived in Table G - MiniMax states its crypto P(reach) three incompatible ways (Table G 0.03-0.08, prose 6-14%, empirical rolling-window histogram 8-18%). Resolved to 0.03-0.08, the table value the headline is built on and the most conservative of the three. Its midpoint therefore carries less weight than its rank-1 position suggests, which is one of the two reasons the universe-wide central (0.03) sits below this row's central (0.05).",
"table_c_feasibility_verdict": "Not graded in Table C; ranked 1st in Table G on P(reach) but carries negative after-tax EV",
"p_ruin_band": "0.40-0.60",
"p_ruin_definition": "Section 12 reports P(ruin) only as a band, never as a row-level central, so p_ruin is null rather than a midpoint. 'Ruin' throughout means terminal wealth <= USD 25 (an experiment-killing loss), NOT literal total loss.",
"table_g_rank": 1
}
},
{
"name": "High-beta single-name equity, full $100 position, 60-90 day hold, unlevered",
"vehicle": "Fractional equity - single name",
"evidence_tier": "T1 cross-sectional volatility / T6 name-specific probabilities",
"citations": [],
"documented_effect_size": "Not expressed as a documented anomaly effect size; rests on single-name idiosyncratic variance",
"post_publication_decay": "n/a - not a published anomaly",
"min_capital_usd": 1.0,
"p_reach_target": 0.04,
"p_ruin": null,
"median_days_to_target": null,
"feasible": false,
"parameters": {
"total_friction_drag_pct_of_stake": "0.02-2.0%",
"after_tax_ev_delta_usd": "-10 to -20",
"p_reach_band": "0.02-0.06",
"note": "A concentration bet whose outcome is driven by the single name's idiosyncratic variance rather than by any documented premium. Under Dubins-Savage this is variance purchase, which is the only mechanism that reaches a fixed target in a subfair game under a deadline.",
"table_c_feasibility_verdict": "Not graded in Table C; ranked 2nd in Table G, negative after-tax EV",
"p_ruin_band": "0.35-0.55",
"p_ruin_definition": "Section 12 reports P(ruin) only as a band, never as a row-level central, so p_ruin is null rather than a midpoint. 'Ruin' throughout means terminal wealth <= USD 25 (an experiment-killing loss), NOT literal total loss.",
"table_g_rank": 2
}
},
{
"name": "CFTC event-contract longshot lottery play (YES at $0.05-$0.10, multi-contract)",
"vehicle": "Kalshi event contracts - longshot tail",
"evidence_tier": "T1 (favorite-longshot bias operates AGAINST the buyer) / T5 fee formula / T6 probabilities",
"citations": [
{
"title": "Explaining the Favorite-Longshot Bias (Snowberg & Wolfers 2010), JPE 118(4), 723-746 - DOI disputed, not asserted",
"doi_or_url": null,
"year": 2010
}
],
"documented_effect_size": "Negative: longshots are overpriced by ~25-30% in the transferred racetrack literature, so the bias runs against the buyer at this end of the curve",
"post_publication_decay": "n/a",
"min_capital_usd": 0.01,
"p_reach_target": 0.005,
"p_ruin": null,
"median_days_to_target": null,
"feasible": false,
"parameters": {
"total_friction_drag_pct_of_stake": "5-25%",
"after_tax_ev_delta_usd": "-20 to -35",
"p_reach_band": "< 0.01",
"fee_note": "At P=0.05 the fee alone consumes 13.3% of stake round trip. Longshots are the expensive regime, not the cheap one - the fee per dollar of stake is 0.07 x (1-P) per side, monotonically decreasing in P.",
"legal_condition": "Same contested Massachusetts access as the other event-contract rows.",
"table_c_feasibility_verdict": "Not graded in Table C; ranked last in Table G with the second-worst P(ruin)",
"p_ruin_band": "0.70-0.90",
"p_ruin_definition": "Section 12 reports P(ruin) only as a band, never as a row-level central, so p_ruin is null rather than a midpoint. 'Ruin' throughout means terminal wealth <= USD 25 (an experiment-killing loss), NOT literal total loss.",
"table_g_rank": 8
}
}
],
"ineffective_strategies": [
{
"name": "Technical-analysis pattern rules (moving averages, RSI, MACD, head-and-shoulders, candlestick reversals)",
"reason": "No signal after 1988. Rules survived data-snooping correction on pre-1988 Dow data (Brock-Lakonishok-LeBaron, confirmed by Sullivan-Timmermann-White's 7,846-rule Reality Check), and the effect vanished thereafter, competed away by institutional algorithmic arbitrage. Under a false-discovery-rate correction on ~100 years of daily Dow data, ZERO rules generate significant out-of-sample excess returns after 5-10 bps of transaction costs. The two source reports disagree on the rule-universe size (15,000+ vs 5,580) and supply incompatible DOIs, so the count is not asserted; the finding both agree on is.",
"evidence_tier": "T1",
"citation": "Park & Irwin (2007), 'What Do We Know About the Profitability of Technical Analysis?', Journal of Economic Surveys 21(4), 786-826, DOI 10.1111/j.1467-6419.2007.00519.x; Bajgrowicz & Scaillet (2012), Journal of Financial Economics 106(3), 473-491, DOI 10.1016/j.jfineco.2012.06.002"
},
{
"name": "Retail day trading (any frequency)",
"reason": "Negative expected value for the median participant, with the largest and cleanest evidence base in the section. Taiwan population study (~1.4M accounts, 15 years): fewer than 1% show predictable, persistent profitability net of fees. Brazil (19,642 equity-futures day traders persisting 300+ days): 97% lost money, only 1.1% earned more than the minimum wage (~USD 54/day), only 0.1% more than USD 300/day. Barber & Odean (2000): the average of 66,465 US households UNDERPERFORMED the market (~16.4% vs 17.9%), with the highest-turnover quintile trailing by ~6.5 points. The mechanism is disposition bias compounded by paying bid-ask spreads to institutional market makers on every round trip - a trader right 50% of the time still loses at a rate set by spread x turnover. Note the widely repeated '99% unprofitable' and '99.9% lose' renderings overstate the papers, whose finding concerns the fraction exhibiting REPEATABLE SKILL.",
"evidence_tier": "T1",
"citation": "Barber, Lee, Liu & Odean (2014), Taiwan day-trading population study; Chague, De-Losso & Giovannetti (2020); Barber & Odean (2000), 'Trading Is Hazardous to Your Wealth', Journal of Finance 55(2), 773-806, DOI 10.1111/0022-1082.00223; Barber, Lee, Liu & Odean (2009), RFS 22(2), 609-632, DOI 10.1093/rfs/hhn046"
},
{
"name": "Leveraged and inverse ETFs held beyond one day",
"reason": "Deterministic path-dependent volatility drag. X_t = X_0 (S_t/S_0)^L exp(0.5(L - L^2) sigma^2 t); the exponential term is negative for every L outside [0,1], which includes every leveraged long and every inverse fund. At L=3 on a FLAT index with sigma = 25% annualized the fund loses exp(-3 x 0.25^2 x 0.25) - 1 = -4.6% over 90 days purely from path volatility. IMPORTANT CORRECTION carried by the merge and stated by no single source: leverage GENUINELY RAISES the probability of hitting a fixed doubling target (a 2x fund needs +41% on the underlying, not +100%) while LOWERING expected value. Those are different quantities and the drag does not close the gap. The instrument is a variance purchase, which under Dubins-Savage is not automatically irrational in a subfair fixed-target game. What kills it is the combination: the drag is deterministic and always adverse, an expense ratio sits on top, and the same variance is available through instruments with bounded downside and no daily-reset penalty.",
"evidence_tier": "T1",
"citation": "Avellaneda & Zhang (2010), SIAM Journal on Financial Mathematics, DOI 10.1137/090771333; Cheng & Madhavan (2009) [T4]; SEC investor bulletin [T5]. Trainor (2010)'s claimed 25-75% annual underperformance was DROPPED as irreconcilable by an order of magnitude with the shared formula."
},
{
"name": "Penny stocks, OTC securities, and pink sheets",
"reason": "Structurally negative expected value driven by spread capture and manipulation rather than by directional risk - the failure mode is a transfer, not a statistical one. Bid-ask spreads run 10% to 50% of share price across 1,000+ microcap and OTC issues, with toxic convertible death-spiral dilution, pervasive pump-and-dump, and long-term returns approaching -100%. A 10% spread means a position must appreciate 11% to break even on a round trip; a 50% spread means it must double simply to exit at cost. The upside case - entering a pump early and exiting before the dump - is the documented mechanism by which retail LOSES in this venue: coordinated operators control the timing and retail flow is the exit liquidity. The distribution is adverse in expectation AND adverse conditional on the scenario the buyer is hoping for.",
"evidence_tier": "T1",
"citation": "Bradley, Cooney, Dolvin & Jordan (2014), 'Penny Stock IPOs', Journal of Banking & Finance 43, 62-73, DOI 10.1016/j.jbankfin.2014.03.003. MiniMax's competing 7-12% pink-sheet markup (Li & Zheng 2020) was dropped as unplaceable and possibly fabricated."
},
{
"name": "Social-media signals, meme momentum, and sentiment-only strategies",
"reason": "No replicated positive expected value. Social sentiment metrics LAG price action - retail buys at peak sentiment precisely as institutional shorting and mean reversion begin. The signal is not absent; it is late. Two deeper reasons it cannot be validated even in principle: (1) the entire meme-equity literature is contingent on one event window in January 2021 on a specific set of retail platforms, and there has been no comparable second attention shock, so there is no out-of-sample replication and a strategy with one observation cannot be validated at any confidence level; (2) the sentiment feature space is effectively unbounded - text polarity, emoji counts, hashtag frequency, follower counts, retweet velocity, subreddit post volume, and arbitrary combinations and lags of each - so under any honest multiple-testing correction the expected value of the best-performing discovered signal converges to zero. This is why sentiment strategies backtest well and trade badly.",
"evidence_tier": "T2",
"citation": "Nofsinger, Sault & Shank (2021), Journal of Behavioral Finance 22(4), 412-428, DOI 10.1080/15427560.2021.1963232; Da, Engelberg & Gao (2011), 'In Search of Attention', Journal of Finance, DOI 10.1111/j.1540-6261.2010.01629.x (direction retained, its 0.22%-per-SD point estimate downgraded to T6); Pedersen (2022) on the GameStop episode [T2], venue flagged for verification"
},
{
"name": "Naive machine learning on price series without purged cross-validation",
"reason": "Guaranteed to overstate out-of-sample performance by four compounding mechanisms. (1) LABEL OVERLAP: forward-return labels over horizon h mean the label at t is a function of prices in [t, t+h], so adjacent observations share outcome information and standard k-fold cross-validation's exchangeability assumption fails - the model memorizes a shared outcome rather than learning a predictive relationship. Characteristic signature: in-sample Sharpe above 4.0 collapsing to zero or negative live. (2) SERIAL CORRELATION inflates the effective sample, so nominal significance thresholds computed on row counts are systematically too permissive. (3) NON-STATIONARITY means a model fitted to one regime estimates a transient artifact rather than a causal law, and markets uniquely react to being modeled. (4) TRIAL-COUNT INFLATION: every architecture, feature set, lookback window and hyperparameter grid is a trial and the reported Sharpe is the maximum over trials; the Probability of Backtest Overfitting exceeds 50% at trial counts a single practitioner reaches in an afternoon. The failure is not that ML does not work on markets - it is that the standard validation toolchain is invalid on this data class.",
"evidence_tier": "T1",
"citation": "Lopez de Prado (2018), Advances in Financial Machine Learning, Wiley [T4]; Cont (2001), Quantitative Finance, DOI 10.1080/713665670 [T1]; Bailey, Borwein, Lopez de Prado & Zhu (2014) on PBO and the Deflated Sharpe Ratio [T1]; Hou, Xue & Zhang (2020), RFS 33(5), 2019-2133, DOI 10.1093/rfs/hhy131; McLean & Pontiff (2016), Journal of Finance 71(1), 5-32, DOI 10.1111/jofi.12365"
},
{
"name": "Copy-trading, signal services, and paid subscription systems",
"reason": "Negative expected value for the subscriber, argued from equilibrium and from the fund-persistence literature by analogy. This is the weakest-evidenced category in the section and the merged report says so rather than manufacturing support - NO source supplies a usable T1 or T2 citation bearing directly on retail copy-trading. THE EQUILIBRIUM ARGUMENT, which needs no citation: if the provider has genuine skill, the profit-maximizing deployment of that skill is proprietary capital and any published signal is a marketing artifact; if the provider lacks skill, the signal is noise sold at a price; in the intermediate case, subscriber flow degrades the very signal being sold because subscribers execute after the provider and into the price impact the aggregate subscription creates. THE MICROSTRUCTURE ARGUMENT: adverse selection, execution latency, and provider-first execution mean any signal with genuine short-horizon content is worth less to the subscriber than to the provider by exactly the latency - and short-horizon content is what these services predominantly sell. THE REPORTING BIAS: published track records are gross of fees, uncorrected for multiple testing across the platform's provider population, and survivorship-affected. At USD 100 a USD 20 monthly subscription consumes 60% of the stake over the window before a single trade.",
"evidence_tier": "T1 by analogy; T4 direct",
"citation": "Carhart (1997), Journal of Finance, DOI 10.1111/j.1540-6261.1997.tb03808.x; Fama & French (2010), 'Luck Versus Skill in the Cross-Section of Mutual Fund Returns', Journal of Finance - NO DOI supplied by any source report and none invented. All direct copy-trading citations offered by the source reports were dropped as topic labels, unverifiable, or fabricated."
},
{
"name": "Martingale, anti-martingale, and progressive position sizing",
"reason": "Ruin in finite time under a finite bankroll, with ruin probability rising toward certainty in the trade count. ABSORPTION PROOF: with B0 = USD 100 and b1 = USD 1, bet k+1 after k consecutive losses is 2^k b1 and cumulative loss is (2^k - 1) b1. The account funds six bets (cumulative USD 63) and CANNOT fund the seventh (USD 64 required against USD 37 remaining). RUIN PROBABILITY at a FAIR p = 0.5: approximately 1 - exp(-N x 0.5 x 0.5^7), giving 54.2% over N = 200 trades and 85.8% over N = 500, converging to 1.00 as N grows. Note the input - the scheme does not require an unfavorable edge to destroy the account, only enough repetitions; realistic transaction costs push p below 0.5 and accelerate every figure. WHY THEORY SAYS DO NOT PLAY: for a game with no edge or a negative edge the Kelly fraction is zero or negative - the optimal bet size is nothing. Martingale is neither bold nor Kelly; it is a timid-play schedule with an exploding tail, the worst available combination for a fixed-target problem, because it maximizes the number of trials (and therefore cumulative ruin hazard) while never concentrating enough stake into any single trial to move the target-hitting probability. Anti-martingale is not ruinous in the same finite-time sense but maximizes exposure at the point of maximum accumulated gain. The general principle: position sizing can amplify a positive edge but cannot manufacture one from a negative expectation - every sizing rule is a linear operator on the per-trade expectation and none changes its sign.",
"evidence_tier": "T1",
"citation": "Kelly (1956), Bell System Technical Journal 35(4), 917-926, DOI 10.1002/j.1538-7305.1956.tb03809.x; Dubins & Savage (1965), How to Gamble If You Must: Inequalities for Stochastic Processes, ISBN 978-0486780641"
},
{
"name": "Overfit backtests as a category",
"reason": "The meta-cause of false confidence in every other category. The mechanism is not sloppiness - it is that the search procedure that finds a strategy is also the procedure that inflates its apparent performance, and the inflation is invisible from inside the search. Minimum Backtest Length: MBL > (2 ln N)/E[SR]^2 x (1 - g1 E[SR] + ((g2-1)/4) E[SR]^2). Testing N = 100 strategy variations on three years of daily data mathematically guarantees an in-sample Sharpe above 2.0 by chance alone, and preventing that false discovery at N = 100 requires more than twelve years of data. (Gemini's own JSON weakened this to N >= 20; the body figure of N = 100 is carried and the internal discrepancy recorded - the point survives either number.) WHY IT GENERATES CONFIDENCE RATHER THAN DOUBT: each individual decision - trying a second lookback window, dropping a delisted name because its data is messy, using the vendor's adjusted price series - is locally reasonable and none announces itself as a trial, so the researcher's subjective count of hypotheses tested is systematically far below the true N and even a researcher who intends to apply a multiple-testing correction applies it at the wrong N. The result is a backtest whose apparent quality rises monotonically with effort, which is precisely the feedback signal a diligent person will pursue.",
"evidence_tier": "T1",
"citation": "Bailey, Borwein, Lopez de Prado & Zhu (2014), 'Pseudo-Mathematics and Financial Charlatanism', Notices of the AMS 61(5), 458-471, DOI 10.1090/noti1105; Harvey, Liu & Zhu (2016), RFS 29(1), 5-68, DOI 10.1093/rfs/hhv059"
}
],
"imported_techniques": [
{
"name": "Brier score",
"source_domain": "Meteorology",
"citation": "Brier (1950), Monthly Weather Review 78(1): 1-3, doi:10.1175/1520-0493(1950)078<0001:VOFEIT>2.0.CO;2 [T1] (Qwen, Gemini, MiniMax). One source attached an ACL 2024 NLP paper as the URL; dropped as fabricated, citation retained.",
"transfer_mechanism": "Score subjective probabilities against binary event-contract settlements; aggregate over a rolling window of N >= 50 contracts BEFORE risking capital. A binary event contract is a probability forecast with a cash settlement attached, so transfer is essentially free of conceptual adaptation cost.",
"transfer_risk": "Assumes the forecaster cannot move the outcome, and treats market prices as rational so systematic biases are absorbed as noise. Endogenous once size moves the book.",
"python_implementation": "scores; properscoring.brier_score; or ~5 lines of NumPy"
},
{
"name": "Brier skill score (BSS)",
"source_domain": "Meteorology",
"citation": "Brier (1950); formalized in Murphy (1973), Monthly Weather Review 101(7): 603-608, doi:10.1175/1520-0493(1973)101<0603:HATMOT>2.0.CO;2 [T1]",
"transfer_mechanism": "Normalize the Brier score against a reference climatology to express skill relative to a naive forecast.",
"transfer_risk": null,
"python_implementation": "Manual; Brier plus a reference baseline"
},
{
"name": "Murphy score decomposition, BS = REL - RES + UNC",
"source_domain": "Meteorology",
"citation": "Brier (1950) / Gneiting & Raftery (2007), JASA 102(477): 359-378, doi:10.1198/016214506000001437 [T1]",
"transfer_mechanism": "Separate reliability, resolution and irreducible event uncertainty in event-contract pricing, isolating which component the trader actually controls. This is the decomposition finance most often omits, and it is the formal machinery behind 'realized calibration' - the one output of a ~63-trade sample that is genuinely well-powered, because calibration is estimable from far fewer observations than edge is.",
"transfer_risk": "The atmosphere is non-adversarial; markets have adversarial feedback loops.",
"python_implementation": "scores; manual partition"
},
{
"name": "Logarithmic score (log-loss)",
"source_domain": "Meteorology / information theory",
"citation": "Good (1952), JRSS B 14(1): 107-114, doi:10.1111/j.2517-6161.1952.tb00085.x [T1]",
"transfer_mechanism": "Score subjective binary probabilities with a proper rule that penalizes hedging harder than Brier.",
"transfer_risk": "Near-infinite penalty when p-hat approaches 0 and the outcome occurs; one mis-stated near-certainty dominates the entire record.",
"python_implementation": "sklearn.metrics.log_loss; NumPy"
},
{
"name": "Continuous Ranked Probability Score (CRPS)",
"source_domain": "Meteorology",
"citation": "Matheson & Winkler (1976), Management Science 22(10): 1087-1096, doi:10.1287/mnsc.22.10.1087; decomposition in Hersbach (2000), Weather & Forecasting 15(5): 559-570 [T1]",
"transfer_mechanism": "Score distributional forecasts for multi-outcome events such as a posterior over the FOMC rate path. It is the only scoring rule that does not require binning choices that alter the score.",
"transfer_risk": "The fitted predictive distribution may not match the actual, and the reference climatology is non-stationary.",
"python_implementation": "properscoring.crps_empirical / crps_gaussian; scores; scoringrules"
},
{
"name": "Sharpness-calibration decomposition",
"source_domain": "Meteorology",
"citation": "Gneiting, Balabdaoui & Raftery (2007), JRSS B 69(2): 243-268, doi:10.1111/j.1467-9868.2007.00543.x [T1]",
"transfer_mechanism": "Decompose CRPS into reliability plus sharpness to diagnose a forecaster who is calibrated but uninformative. Maximum achievable sharpness given calibration is a property of the data-generating process, not of the model class.",
"transfer_risk": "The true distribution is unobservable and non-stationary, so the decomposition's reference drifts.",
"python_implementation": "Manual computation on CRPS sub-components"
},
{
"name": "Proper scoring rule convention",
"source_domain": "Meteorology",
"citation": "Gneiting & Raftery (2007), JASA 102(477): 359-378, doi:10.1198/016214506000001437 [T1]. NOTE: one source's markdown gives 10.1198/... while its own JSON appendix gives 10.1188/...; the 10.1198 form is 2-of-3 and internally consistent across two independent reports.",
"transfer_mechanism": "Constrain in-strategy loss functions to the class that is minimized in expectation by the true conditional distribution, which induces honest probability reports. Improper scores reward hedging.",
"transfer_risk": "No external scorer enforces propriety on a solo trader; the discipline is internal only.",
"python_implementation": "Implement scoring rules directly"
},
{
"name": "Reliability diagram with bootstrap confidence bands",
"source_domain": "Meteorology",
"citation": "Broecker & Smith (2007), Weather & Forecasting 22(3): 651-661, doi:10.1175/WAF993.1 [T1]",
"transfer_mechanism": "Visual test of whether the calibration curve lies inside the 95% band of the diagonal before further capital is deployed.",
"transfer_risk": "Assumes forecasts are exchangeable across time and forecaster; real forecasts are autocorrelated and the forecaster evolves. In-sample fit risk.",
"python_implementation": "uncertainty-toolbox (unmaintained); matplotlib plus bootstrap bands"
},
{
"name": "PIT / rank histogram",
"source_domain": "Meteorology",
"citation": "Dawid (1984), Z. Wahrsch. verw. Gebiete 60: 305-313, doi:10.1007/BF00524500; Hamill (2001), Monthly Weather Review 129(3): 550-560 [T1]",
"transfer_mechanism": "Diagnose distributional calibration for non-binary events such as rate paths or index levels.",
"transfer_risk": "Same exchangeability violation as the reliability diagram.",
"python_implementation": "Manual histogram plus a KS test"
},
{
"name": "Ensemble forecasting",
"source_domain": "Meteorology",
"citation": "Leith (1974), Monthly Weather Review 102(6): 409-418, doi:10.1175/1520-0493(1974)102<0409:TSOMCF>2.0.CO;2 [T1]. One source attributed it to a 2011 textbook with a URL pointing at a WMO magazine article about a 2024 prize lecture; dropped as falsified and as a textbook rather than an originating work.",
"transfer_mechanism": "Bootstrap N draws of the signal distribution to quantify output uncertainty and tail risk - e.g. 1,000 simulated paths for a CPI release.",
"transfer_risk": "Assumes a physics-consistent multi-member ensemble that a retail participant does not possess and must approximate by bootstrapping; the bootstrap sampling distribution may not match true uncertainty, and financial feedback loops correlate errors across members.",
"python_implementation": "numpy.random bootstrap; conformal bands as a sanity check"
},
{
"name": "Model Output Statistics (MOS)",
"source_domain": "Meteorology",
"citation": "Glahn & Lowry (1972), J. Applied Meteorology 11(8): 1203-1211 [T1]",
"transfer_mechanism": "Regress raw model output onto historically observed venue mid-prices to remove systematic bias. A strict improvement whenever the raw output is systematically biased.",
"transfer_risk": "Assumes a stable bias-to-surface mapping; recalibration warps over months and requires rolling exponentially-weighted recomputation.",
"python_implementation": "scipy.optimize.minimize; constrained linear regression"
},
{
"name": "Base-rate outside-view priming",
"source_domain": "Judgmental forecasting (Good Judgment Project)",
"citation": "Kahneman & Tversky (1973), Psychological Review 80(3): 237-251, doi:10.1037/h0034749 [T1]. One source attributed it to a 2015 trade book with a 2010 newsletter URL that predates it.",
"transfer_mechanism": "Force explicit reference-class construction before any probability claim is entered in the ledger.",
"transfer_risk": "Reference-class composition drifts and its selection is arbitrary enough to introduce bias.",
"python_implementation": "Elicitation-UI discipline; no library required"
},
{
"name": "Extremizing - linear blend, alpha in [0.05, 0.15] per iteration",
"source_domain": "Judgmental forecasting (Good Judgment Project)",
"citation": "Baron, Mellers, Tetlock, Stone & Ungar (2014), cited as Psychological Science 25(2): 437-444, doi:10.1177/0956797613504262 [T2] - TITLE AND VENUE FLAGGED BY THE CITING SOURCE'S OWN DIGEST as describing a paper about willingness to forecast rather than about extremizing. Reproduced exactly as given, with the mismatch recorded for downstream verification; no replacement citation was invented.",
"transfer_mechanism": "Shift the private forecast p-hat away from a market consensus m toward 0 or 1 by a small fraction per update, under a Brier objective.",
"transfer_risk": "The choice of alpha is fragile to regime and the score is asymmetric under partial pooling. Requires holdout validation before production use.",
"python_implementation": "Manual linear combination"
},
{
"name": "Extremizing - logit form, scaling exponent d = 1.4",
"source_domain": "Judgmental forecasting (Good Judgment Project)",
"citation": "Satopaa et al. (2014), Annals of Applied Statistics 8(2): 916-940, doi:10.1214/14-AOAS752; Baron et al. (2014), Decision Analysis 11(2): 133-145, doi:10.1287/deca.2014.0293 [T2]",
"transfer_mechanism": "Logit-scale an underconfident crowd or model-ensemble consensus before pricing it against a contract. NOT COMPARABLE to the linear-blend alpha above: these parameterize different operations on different scales, and printing them adjacent invites a comparison that does not exist. Neither adjudicates the other.",
"transfer_risk": "GJP questions were static and long-horizon; order books reprice instantly on news.",
"python_implementation": "scipy.optimize; NumPy"
},
{
"name": "Trimmed-mean / geometric-mean aggregation",
"source_domain": "Judgmental forecasting (Good Judgment Project)",
"citation": "Mellers, Stone, Murray et al. (2015), Perspectives on Psychological Science 10(3): 267-281, doi:10.1177/1745691615576804 [T1]",
"transfer_mechanism": "Aggregate several probability sources by trimmed mean rather than simple average; Brier-improving under proper scoring.",
"transfer_risk": "Forecasters are non-independent and biased in different directions; bias-correct BEFORE aggregating, not after.",
"python_implementation": "numpy.mean on a trimmed array"
},
{
"name": "Linear opinion pooling / combining forecasts",
"source_domain": "Judgmental forecasting (Good Judgment Project)",
"citation": "Clemen (1989), International Journal of Forecasting 5(4): 559-583, doi:10.1016/0169-2070(89)90012-8 [T1]; Cooke (1981), Experts in Uncertainty [T2]",
"transfer_mechanism": "Pool venue consensus, economist-survey medians (e.g. the Survey of Professional Forecasters) and private signals into one probability.",
"transfer_risk": null,
"python_implementation": "NumPy weighted combination"
},
{
"name": "Track-record / accuracy-weighted aggregation",
"source_domain": "Judgmental forecasting (Good Judgment Project)",
"citation": "Satopaa (2014), PhD thesis [T3] - institutional handle flagged as implausible by the citing source's own digest; operational form in Mellers et al. (2015)",
"transfer_mechanism": "Weight each probability source by its own historical accuracy rather than equally.",
"transfer_risk": "Track-record validity under non-stationary environments is explicitly an open question in the source itself.",
"python_implementation": "Manual weighting; see the IRT rows for the formal treatment"
},
{
"name": "Frequent updating as tournament discipline",
"source_domain": "Judgmental forecasting (Good Judgment Project)",
"citation": "Mellers et al. (2015), doi:10.1177/1745691615576804 [T1]",
"transfer_mechanism": "Treat each macro release as a scored 'tournament tick' and re-price immediately rather than holding a stale position. Frequent updating - not raw cognitive talent - is the dominant behaviorally measurable contributor to forecasting score.",
"transfer_risk": "CALIBRATION COLLAPSE UNDER STAKE SIZE: probability estimates compress systematically toward 0.5 when the bid-ask spread is non-trivial, a direct violation of proper-scoring theory. GJP ran with low-stakes incentives, publicly observable ground truth, and skill-selected participants; this problem inverts all three.",
"python_implementation": "Event-driven loop; no library required"
},
{
"name": "Prediction markets versus prediction polls",
"source_domain": "Judgmental forecasting (Good Judgment Project)",
"citation": "Atanasov, Reshetar, Zhang & Zwick (2020), cited as Management Science 66(9): 4076-4094, doi:10.1287/mnsc.2019.2269 [T2] - AUTHOR LIST, VOLUME AND DOI FLAGGED AS IMPLAUSIBLE for the title given, by the citing source's own digest.",
"transfer_mechanism": "Use venue-implied probabilities as an INPUT to the trader's hierarchical pool rather than as a competitor to it.",
"transfer_risk": "Endogeneity at size. The directional claim - do superforecasters beat markets? - is contested between sources and is EXCLUDED: one asserts they consistently outperform prediction markets (sourced to a vendor-interested self-published PDF), the other says performance converges with market-implied probabilities and states the direction two different ways internally. Only the uncontested component survives: structured aggregation of trained forecasters outperforms unstructured individual judgment.",
"python_implementation": "Venue API extraction; NumPy"
},
{
"name": "Delphi method",
"source_domain": "Judgmental forecasting (Good Judgment Project)",
"citation": "Rowe & Wright (1999) [T2] - URL is a course-site mirror, not a publisher host",
"transfer_mechanism": "Structured multi-round elicitation and aggregation of expert judgments into a consensus forecast.",
"transfer_risk": "Groupthink and facilitator influence; less applicable to anonymous online markets, and a solo retail participant cannot run it at all.",
"python_implementation": "Not applicable to a single participant"
},
{
"name": "Verbal-to-numeric elicitation for rare events",
"source_domain": "Judgmental forecasting (Good Judgment Project)",
"citation": "Fischhoff & Davis (2014), WIREs Climate Change, doi:10.1002/wcc.318 [T2] - TOPIC MISMATCH FLAGGED: the cited paper is on climate-uncertainty communication.",
"transfer_mechanism": "Convert qualitative conviction into a reference-class PMF before it enters the ledger.",
"transfer_risk": null,
"python_implementation": "Elicitation UI"
},
{
"name": "Systematic review of superforecasting research",
"source_domain": "Judgmental forecasting (Good Judgment Project)",
"citation": "Himmelstein & Stahl (2023), Judgment and Decision Making 18: e22, doi:10.1017/jdm.2023.23 [T2]",
"transfer_mechanism": "Provenance check on every GJP-derived technique before adoption. NOTE: one source asserts 'at least three independent meta-analyses from 2018-2023 reproduce the core finding' but names only this one; treat the replication claim as supported by ONE named systematic review, not three.",
"transfer_risk": null,
"python_implementation": "n/a"
},
{
"name": "Cramer-Lundberg ruin model and the Lundberg adjustment coefficient",
"source_domain": "Actuarial science",
"citation": "Lundberg (1903); Cramer (1930); modern treatment Asmussen & Albrecher (2010), Ruin Probabilities [T2]. Surplus process U(t) = u + ct - S(t) with ruin probability psi(u) = P(inf U(t) < 0); Lundberg's inequality psi(u) <= exp(-Ru), where R is the unique positive root of lambda + cR = lambda M_X(R).",
"transfer_mechanism": "THE DEEPEST CONCEPTUAL IMPORT IN THE INVENTORY, and the one all three reports reach independently: the doubling problem is literally the DUAL of the actuarial problem - minimize P(ruin) on the path to a target given a fixed maximum loss budget, instead of minimizing P(ruin) on the path to insolvency given a fixed premium stream. Model daily P&L as a surplus process and compute P(the USD 100 stake is depleted before day 90) under the candidate strategy's empirical return distribution; the Lundberg coefficient resolves whether a positive expected log-return is sufficient to make P(ruin) < 1. NOTE: one source printed the adjustment-coefficient condition as 'E[exp(gamma X)] < 1 for some gamma > 0', which is unsatisfiable for any positive claim size since exp(gamma X) > 1 pointwise; that form is excluded and falsely attributed, and the correct forms above are carried.",
"transfer_risk": "Claim sizes are assumed i.i.d. and exogenous to the insurer's activity; trading returns are serially correlated through overnight gaps and macro cycles, non-stationary, and exhibit tail clustering. More subtly, ruin theory does not condition on the data-generating process changing IN RESPONSE TO the analyst's signal, so where the signal correlates with the future evolution of the distribution, ruin estimates are systematically OPTIMISTIC.",
"python_implementation": "numpy.random Monte Carlo; Lundberg exponent in ~50 lines of scipy.optimize; lifelib primitives"
},
{
"name": "Collective risk model - frequency-severity decomposition",
"source_domain": "Actuarial science",
"citation": "Panjer (1981), ASTIN Bulletin 12(1): 22-26, doi:10.1017/S0515036100006615 [T1]",
"transfer_mechanism": "The correct FIRST decomposition of any candidate return-generating process: severity is the right tail of the log-return distribution, where kurtosis dominates and the log-normal right tail is far too thin; frequency is the purged, cross-validated effective observation count per quarter.",
"transfer_risk": "Empirical distributions overfit against the assumed Poisson/negative-binomial family; use an empirical bootstrap as a cross-check. (A kurtosis figure of ~10-20 for daily log-returns of liquid US equities is stated by one source with no source, sample period or universe definition.)",
"python_implementation": "Manual, ~30 lines"
},
{
"name": "Panjer recursion",
"source_domain": "Actuarial science",
"citation": "Panjer (1981), doi:10.1017/S0515036100006615 [T1]",
"transfer_mechanism": "Compute the EXACT finite-horizon aggregate P&L distribution S = X1 + ... + XN for small trade counts (one source specifies K = 10 per quarter), avoiding asymptotic approximations that are worthless at N = 10. That exactness matters here precisely because the sample is tiny.",
"transfer_risk": "Distribution-family mismatch between the assumed compound family and realized returns.",
"python_implementation": "Manual, ~30 lines"
},
{
"name": "Buhlmann credibility, Z = n/(n+K)",
"source_domain": "Actuarial science",
"citation": "Buhlmann (1967), ASTIN Bulletin 4(3): 199-207, doi:10.1017/S0515036100008832 [T1]",
"transfer_mechanism": "Answers the question a backtest cannot: what is the prior probability that a strategy has real edge, given that it appears in the literature at all? Shrink a strategy-edge estimate toward the population mean of pre-registered retail strategies, with weight rising in observation count. AT N = 10 TO 30 TRADES, Z IS SMALL AND THE SHRINKAGE IS SEVERE - which is the correct behavior and also the reason a 90-day live result cannot escape its prior.",
"transfer_risk": "Risk classes are assumed mutually independent; candidate strategies are correlated through shared macro factors. K is unknown for a new strategy.",
"python_implementation": "Manual; PyMC for the hierarchical extension"
},
{
"name": "Buhlmann-Straub credibility",
"source_domain": "Actuarial science",
"citation": "Buhlmann & Straub (1970), Mitt. Ver. Schweiz. Versicherungsmathematiker 70: 111-133 [T1]. One source attributes it to a 2005 textbook, which is a later work and not the originating citation.",
"transfer_mechanism": "Multi-level credibility with an explicit measurement-error structure; blends backtest alpha with retail base rates.",
"transfer_risk": "Assumes the underlying risk process is stationary, which is often false in financial markets; asset returns additionally exhibit tail clustering.",
"python_implementation": "NumPy custom module"
},
{
"name": "Bayesian credibility (credibility as conjugate Bayes)",
"source_domain": "Actuarial science",
"citation": "Jewell (1974), Geneva Papers on Risk and Insurance Theory 1(1): 77-80, doi:10.1007/BF02553258 [T2] - title garbled in the source ('Bayesian Bayesian') and the year/volume pairing flagged.",
"transfer_mechanism": "Establishes Buhlmann credibility as exact Bayes under a conjugate prior, licensing direct Bayesian implementation.",
"transfer_risk": null,
"python_implementation": "scipy.stats conjugate updates"
},
{
"name": "Extreme-value theory - GPD peaks-over-threshold",
"source_domain": "Actuarial science",
"citation": "Embrechts, Kluppelberg & Mikosch (1997), Modelling Extremal Events [T2]; McNeil, Frey & Embrechts (2015), Quantitative Risk Management [T2]",
"transfer_mechanism": "Fit the empirical peaks-over-threshold Generalized Pareto tail on each candidate strategy's worst 5% of observations and VERIFY GPD FIT BEFORE TRUSTING ANY ESTIMATED SHARPE.",
"transfer_risk": null,
"python_implementation": "scipy.stats.genpareto; manual POT fit"
},
{
"name": "Loss-development triangles / chain-ladder",
"source_domain": "Actuarial science",
"citation": "Mack (1993), ASTIN Bulletin 23(2): 213-225, doi:10.1017/S0515036100009412 [T1]",
"transfer_mechanism": "NONE - EXPLICITLY DISCLAIMED. Listed only because a reader will encounter it in the actuarial literature. This is the only instance in any of the three reports of a technique being named and then correctly excluded, and the disclaimer is preserved as stated.",
"transfer_risk": "The source states: 'Largely irrelevant; do not transfer.' Retained as an explicit exclusion, not a recommendation.",
"python_implementation": "chainladder - not needed for this problem"
},
{
"name": "Bayesian nowcasting under reporting delay",
"source_domain": "Epidemiology / public-health nowcasting",
"citation": "Hohle & an der Heiden (2014), Biometrics 70(4): 993-1002, doi:10.1111/biom.12194; generalized in Gunther et al. (2021), Biometrical Journal 63(8): 1575-1593, doi:10.1002/bimj.202000112 [T1]",
"transfer_mechanism": "Daily-updated estimate of a latent macro variable from sparse, delayed observations between scheduled releases. The macro-release calendar is the direct analogue: CPI, NFP, PCE and FOMC releases arrive at irregular intervals with information leaking between them through Fed speeches, equity returns and survey data.",
"transfer_risk": "The reporting system is assumed exogenous, but MACRO REVISIONS ARE STRATEGIC AND BIDIRECTIONAL, NOT MERELY DELAYED - a revision is a decision made by an agency with its own objectives and calendar, whereas reporting delay in an outbreak is a physical and administrative lag. The noise model must jointly specify measurement error, seasonality and revisions or the intervals are overconfident.",
"python_implementation": "PyMC with an explicit reporting-delay layer; arviz for posterior diagnostics"
},
{
"name": "NobBS Bayesian delay nowcasting",
"source_domain": "Epidemiology / public-health nowcasting",
"citation": "McGough, Johansson, Lipsitch & Menzies (2020), PLOS Computational Biology 16(4): e1007735, doi:10.1371/journal.pcbi.1007735 [T1]",
"transfer_mechanism": "Correct reporting delays and backfill in BLS/GDP/CPI release series to produce a current-state estimate. Complementary to rather than competing with Hohle & an der Heiden: that is the originating Bayesian nowcasting method and this is the widely used implementation.",
"transfer_risk": "Clinical reporting delays are physical; economic data are strategically revised.",
"python_implementation": "PyMC"
},
{
"name": "Reporting-delay decomposition and backfill correction",
"source_domain": "Epidemiology / public-health nowcasting",
"citation": "Hohle & an der Heiden (2014); Gunther et al. (2021) [T1]",
"transfer_mechanism": "Re-estimate each past probability once later information completes, then re-score. Applied to the trader's own ledger, a Monday probability revealed as wrong by Tuesday's release is retro-corrected before it enters the calibration record - THIS EXTRACTS MORE INFORMATION PER SETTLED CONTRACT THAN NAIVE SCORING DOES, which is the epidemiological insight that most directly attacks the small-sample problem.",
"transfer_risk": "Past estimates are not only systematically low but also noisy; combine with bootstrap confidence intervals before acting on the correction.",
"python_implementation": "PyMC; manual re-scoring loop"
},
{
"name": "Mixed-frequency nowcasting (MIDAS)",
"source_domain": "Epidemiology / public-health nowcasting (labelled 'macro-econometrics' by its source)",
"citation": "Giannone, Reichlin & Small (2008), Journal of Monetary Economics 55(4): 665-676, doi:10.1016/j.jmoneco.2008.05.010 [T1] - the paper that imported 'nowcasting' into macroeconomics",
"transfer_mechanism": "Kernel-weighted regression on mixed-frequency observations producing a daily probability surface over questions such as 'will CPI exceed 3.0% YoY at the next release?', priced directly against the corresponding event contract.",
"transfer_risk": "Mixed-frequency weighting is fragile to publication-calendar changes; macro data are conditioned on prior announcements and revisions.",
"python_implementation": "statsmodels; manual kernel-weighted lag regression"
},
{
"name": "Hierarchical Bayesian partial pooling",
"source_domain": "Epidemiology / public-health nowcasting",
"citation": "Gelman & Hill (2007), Data Analysis Using Regression and Multilevel/Hierarchical Models [T2]; Carpenter et al. (2017), J. Statistical Software 76(1), doi:10.18637/jss.v076.i01 [T1]",
"transfer_mechanism": "Pool probabilities across venues (Kalshi, ForecastEx, IBKR) with venue-specific intercepts and a common latent-state loading - strictly better than any single venue WHEN THE VENUES ARE PARTIALLY SEGMENTED. Operationally the same as pooling test-positivity across states.",
"transfer_risk": "Cross-market arbitrage collapses the mispricings the pooling is meant to exploit, so the technique's value is INVERSELY PROPORTIONAL to how integrated the venues are; vendor change and regime shift break the pooling structure.",
"python_implementation": "PyMC; cmdstanpy/Stan; brms"
},
{
"name": "Kalman-filter macro nowcaster (fully specified)",
"source_domain": "Epidemiology / public-health nowcasting / Signal processing and industrial statistics",
"citation": "Kalman (1960), J. Basic Engineering 82(1): 35-45, doi:10.1115/1.3662552 [T1]; Harvey (1989), Forecasting, Structural Time Series Models and the Kalman Filter [T2]",
"transfer_mechanism": "THE MOST DIRECTLY EXECUTABLE SPECIFICATION IN THE INVENTORY, requiring no proprietary data: state = (latent inflation nowcast, latent unemployment nowcast, latent recession probability); observations = (released CPI, released NFP, venue-implied probabilities from Kalshi/CME FedWatch); transition = AR(1) latent drift. The posterior mean is a daily probability surface priced directly against contracts.",
"transfer_risk": "The state transition is assumed linear-Gaussian; financial series are fat-tailed and non-Gaussian observation noise degrades the posterior.",
"python_implementation": "filterpy.kalman (unmaintained); statsmodels.tsa.statespace; ~100 lines of NumPy"
},
{
"name": "CUSUM change-point detection",
"source_domain": "Signal processing and industrial statistics",
"citation": "Page (1954), Biometrika 41(1/2): 100-115 [T1] - DOI CONFLICT: one source gives 10.1093/biomet/41.1-2.100 and the other 10.2307/2333009 for the same paper. Irreconcilable without external lookup; both recorded, and the article citation itself is 2-of-2 agreed and not in doubt.",
"transfer_mechanism": "Run CUSUM on running expected log-return, or equivalently on running P&L normalized by per-trade risk; once the cumulative sum exceeds a threshold tuned via in-control Average Run Length, the strategy is declared drifting. In a non-stationary environment CUSUM is CONSERVATIVE - false alarms too rare - which is the correct direction of error for capital protection.",
"transfer_risk": "In-control and post-change distributions are assumed fixed ex ante; both drift continuously in markets. Page-Lorden optimality is exact only in the parametric case, and ARL INFLATION IS SIGNIFICANT WHEN PARAMETERS ARE ESTIMATED FROM DATA - the false-alarm rate is worse than advertised in exactly the regime where the tool is used. Separately, the trader's own position can CAUSE the drift being detected; CUSUM detects this correctly only if reported P&L includes the position-impact component.",
"python_implementation": "ruptures; statsmodels.stats.diagnostic.breaks_cusumolsresid. NOTE: the API path ruptures.detect.cusum cited by one source DOES NOT EXIST - the package exposes search classes Pelt, Binseg, Window, BottomUp, Dynp with cost functions."
},
{
"name": "GLR-CUSUM",
"source_domain": "Signal processing and industrial statistics",
"citation": "Lorden (1971), Annals of Mathematical Statistics 42(6): 1897-1908, doi:10.1214/aoms/1177693014 [T1]",
"transfer_mechanism": "CUSUM that ESTIMATES the post-change parameter rather than fixing it - the recommended variant precisely because post-degradation behaviour is unknown ex ante.",
"transfer_risk": "Requires an estimate of post-change parameters; the same continuous-drift problem as CUSUM, plus strategy detection bias where the position itself causes the drift being measured.",
"python_implementation": "Manual implementation; ruptures"
},
{
"name": "Bayesian Online Change-Point Detection (BOCPD)",
"source_domain": "Signal processing and industrial statistics",
"citation": "Adams & MacKay (2007), arXiv:0710.3742 [T3]; refereed treatment Fearnhead & Liu (2007), JRSS B 69(4): 589-605, doi:10.1111/j.1467-9868.2007.00545.x [T1]",
"transfer_mechanism": "Returns a posterior over RUN LENGTH, updating in O(N) per step and behaving acceptably at small sample sizes. Feed daily P&L in with a hazard rate tuned to the expected strategy half-life; P(run length > k) is the strategy's instantaneous credibility. Applied one level down, it detects order-book regime shifts and volatility breaks for stop-out triggering.",
"transfer_risk": "The hazard/run-length prior is fragile and concept drift produces multiple overlapping changes; assumes Gaussian white noise where returns are jump-diffusion. Mitigation: a heavy-tailed run-length prior.",
"python_implementation": "ruptures; bayesian-changepoint-detection (existence unverified)"
},
{
"name": "Sequential Probability Ratio Test (SPRT)",
"source_domain": "Signal processing and industrial statistics AND Clinical-trial methodology (dual-domain)",
"citation": "Wald (1945), Annals of Mathematical Statistics 16(2): 117-186, doi:10.1214/aoms/1177731118 [T1]. THE SINGLE TECHNIQUE ALL THREE REPORTS NAME. One source attached an Instagram Reel as the originating URL; dropped as fabricated.",
"transfer_mechanism": "Test H0 (win rate = 50%) against H1 (60%) at alpha = 0.05, beta = 0.20 as a stopping rule for a single strategy. For a true 60% win rate the test terminates on average after ~30 trades ([T6] on the figure - stated without formula, parameters or derivation), while under the null it nearly always runs to its upper bound. THE 30-TRADE FIGURE IS THE POINT: the number of trades required to DISTINGUISH a 60% win rate from a coin flip is roughly the same order as the total number of trades a 90-day USD 100 experiment can execute under T+1 settlement. The experiment sits at the resolution boundary of its own test.",
"transfer_risk": "Assumes i.i.d. observations; trade P&L is serially correlated through overnight gaps and macro cycles. The remedy - compute the effective sample size of the trade sequence and use that in the threshold computation - is sound, but its stated attribution (Bartlett's formula located in an interior section of Wald 1945, with no Bartlett citation anywhere in the file) is not. Binary-hypothesis design is additionally awkward for continuous forecasts.",
"python_implementation": "Manual computation with an ESS correction; a gsDesign port for the interim-monitoring form"
},
{
"name": "Wald-Wolfowitz SPRT optimality",
"source_domain": "Signal processing and industrial statistics",
"citation": "Wald & Wolfowitz (1948), Annals of Mathematical Statistics 19: 326-329, doi:10.1214/aoms/1177699121 [T1]",
"transfer_mechanism": "Establishes that SPRT minimizes expected sample size among all tests at the same alpha and beta - the guarantee that makes SPRT worth using at N ~ 30, and the result that also underwrites group-sequential clinical-trial design.",
"transfer_risk": null,
"python_implementation": "n/a - a theoretical guarantee"
},
{
"name": "Shewhart control chart",
"source_domain": "Signal processing and industrial statistics",
"citation": "Shewhart (1924), Economic Control of Manufactured Product; Montgomery (2019), Introduction to Statistical Quality Control, 8th ed. [T2]",
"transfer_mechanism": "Lightweight three-sigma regime detection on P&L.",
"transfer_risk": "Low statistical power against small shifts - which are the shifts most likely to matter at this sample size.",
"python_implementation": "Manual; dashboard"
},
{
"name": "EWMA control chart",
"source_domain": "Signal processing and industrial statistics",
"citation": "Roberts (1959), Technometrics 1(3): 239-250, doi:10.1080/00401706.1959.10489860 [T1]",
"transfer_mechanism": "THE CHEAPEST INSTRUMENT IN THE INVENTORY: a single smoothed deviation-from-target with two-sigma bands; a breach declares regime change.",
"transfer_risk": "Sensitive to the smoothing-parameter choice, and a regime change produces a permanent shift the chart treats as transient. The low cost buys correspondingly low resolution. Mitigation: multiple horizons plus CUSUM as backup.",
"python_implementation": "Manual, ~10 lines"
},
{
"name": "Extended / Unscented Kalman filter",
"source_domain": "Signal processing and industrial statistics",
"citation": "Julier & Uhlmann (1997), Proc. AeroSense; Julier & Uhlmann (2004), Proc. IEEE 92(3): 401-422, doi:10.1109/JPROC.2004.823170 [T1]",
"transfer_mechanism": "Latent-state estimation where the observation or transition map is nonlinear.",
"transfer_risk": null,
"python_implementation": "filterpy"
},
{
"name": "Particle filter",
"source_domain": "Signal processing and industrial statistics",
"citation": "Gordon, Salmond & Smith (1993), IEE Proc. F 140(2): 107-113, doi:10.1049/ip-f-2.1993.0014 [T1]",
"transfer_mechanism": "Non-Gaussian latent-state estimation - e.g. which of three discrete volatility regimes is active; ~150 lines for a one-dimensional state. The escalation from Kalman to particle filtering is the CORRECT response to jump-diffusion returns rather than an optional refinement.",
"transfer_risk": "Computational cost is the binding constraint; mitigation is conjugate approximation or Rao-Blackwellization where feasible.",
"python_implementation": "filterpy.monte_carlo"
},
{
"name": "Rasch model (1-parameter IRT)",
"source_domain": "Psychometrics / item-response theory",
"citation": "Rasch (1960), Probabilistic Models for Some Intelligence and Attainment Tests [T2]",
"transfer_mechanism": "Treat each historical forecast as an item and the trader's calibration as a single latent-trait parameter, estimating item difficulty jointly so that easy and hard contracts are not scored alike. The output is a calibration estimate that PROPERLY ACCOUNTS FOR THE DIFFICULTY OF THE QUESTIONS FACED, which naive Brier scoring does not.",
"transfer_risk": "THE TRIVIAL-N PROBLEM DOMINATES: at N = 10-30 contracts the estimator is not identified. The trader has enough observations to fit a Bayesian IRT with strong priors, not enough to fit a Rasch model on their own forecasting edge in one quarter. Separately, psychometric traits are approximately stable across a test session whereas trading skill is state-dependent - on stake size, on fatigue, on regime - which violates exchangeability more severely than concept drift in a source does.",
"python_implementation": "pyirt; manual EM"
},
{
"name": "2-parameter logistic IRT",
"source_domain": "Psychometrics / item-response theory",
"citation": "Birnbaum (1968), in Lord & Novick, Statistical Theories of Mental Test Scores [T2]",
"transfer_mechanism": "Estimate each information source's discrimination and difficulty separately rather than as a single accuracy number.",
"transfer_risk": "Same trivial-N and non-stationarity problems as the Rasch model.",
"python_implementation": "pyirt"
},
{
"name": "3-parameter logistic IRT (adds a guessing parameter)",
"source_domain": "Psychometrics / item-response theory",
"citation": "Rasch (1960) / Lord (1980) [T2]/[T3]",
"transfer_mechanism": "Separate forecaster skill theta from contract difficulty b_j AND guessing c_j - the correct structure for binary contracts where a coin flip scores 50%.",
"transfer_risk": "Psychometric traits are assumed stable; trader skill fluctuates with stake, regime and fatigue.",
"python_implementation": "scipy.optimize custom 3PL"
},
{
"name": "Polytomous / graded-response IRT",
"source_domain": "Psychometrics / item-response theory",
"citation": "Samejima (1969), Psychometrika 34(4): 1-97, doi:10.1007/BF03390160 [T2] - pagination flagged as a monograph supplement rather than a regular article",
"transfer_mechanism": "Extends the IRT family to ordered multi-outcome contracts rather than binary settlements.",
"transfer_risk": null,
"python_implementation": "pyirt extensions"
},
{
"name": "Hierarchical / Bayesian IRT",
"source_domain": "Psychometrics / item-response theory",
"citation": "Fox (2010), Bayesian Item Response Modeling [T2]",
"transfer_mechanism": "THE IDENTIFIED ALTERNATIVE at this sample size: strong priors plus population pooling in place of a free Rasch fit.",
"transfer_risk": "Item difficulty drifts; requires time-bounded parameter estimation.",
"python_implementation": "PyMC"
},
{
"name": "Empirical-Bayes (EAP) ability estimation",
"source_domain": "Psychometrics / item-response theory",
"citation": "Bock & Mislevy (1982), Applied Psychological Measurement 6(4): 431-444, doi:10.1177/014662168200600405 [T1] - author initials flagged as transposed in the source",
"transfer_mechanism": "Posterior-mean ability estimate that is stable at small N, unlike maximum likelihood.",
"transfer_risk": null,
"python_implementation": "pyirt; manual EAP quadrature"
},
{
"name": "Hierarchical rater model",
"source_domain": "Psychometrics / item-response theory",
"citation": "Patz, Junker, Johnson & Mariano (2002), ETS Research Report [T5] - grey literature; a peer-reviewed version exists and would be the better citation",
"transfer_mechanism": "Cluster signals by source with source-level and contract-level parameters; MCMC posterior over each source's reliability.",
"transfer_risk": "Rater errors are correlated through shared source bias - this model is itself the stated mitigation for the Dawid-Skene conditional-independence violation. MCMC convergence is the practical risk.",
"python_implementation": "PyMC"
},
{
"name": "Dawid-Skene latent-truth model",
"source_domain": "Psychometrics / item-response theory",
"citation": "Dawid & Skene (1979), JRSS C 28(1): 20-28, doi:10.2307/2346806 [T1]",
"transfer_mechanism": "Treat each candidate signal as a RATER and each contract resolution as an ITEM; EM jointly estimates latent truth and per-source error rates over a sliding window of N = 100 contracts, and the latent-truth estimate becomes the pooled probability. This is the same problem as crediting a prediction source - a forecaster, an indicator, an NLP sentiment model, an FOMC statement - with empirical reliability.",
"transfer_risk": "ASSUMES CONDITIONAL INDEPENDENCE OF RATER ERRORS, which correlated signals (all reading the same news) violate. Non-stationary source quality biases the estimates; remedy is a time-bounded rolling re-fit. Contract resolutions are additionally not exchangeable in difficulty - a 99%-probability Fed contract is procedurally easier than a 51%-probability contested-election contract, and ignoring this concentrates all apparent Brier improvement on easy items.",
"python_implementation": "Manual EM, ~30 lines"
},
{
"name": "Generalizability theory (G-theory)",
"source_domain": "Psychometrics / item-response theory",
"citation": "Cronbach, Gleser, Nanda & Rajaratnam (1972), The Dependability of Behavioral Measurements [T2]",
"transfer_mechanism": "Decompose reliability into within-source, between-source and item-heterogeneity variance - separating 'this strategy is fragile to the choice of source' from 'this strategy's signal quality is genuinely high'.",
"transfer_risk": null,
"python_implementation": "Variance-components estimation; statsmodels mixed models"
},
{
"name": "Shannon entropy and mutual information",
"source_domain": "Information theory",
"citation": "Shannon (1948), Bell System Technical Journal 27(3): 379-423 and 27(4): 623-656, doi:10.1002/j.1538-7305.1948.tb01338.x [T1]",
"transfer_mechanism": "Quantify outcome uncertainty and the dependence between a candidate signal and the realized outcome; RANK SIGNALS BY ESTIMATED MI RATHER THAN BY RAW ACCURACY, which is blind to redundancy - a low-accuracy but high-conditional-MI signal can outperform a high-accuracy but near-redundant one.",
"transfer_risk": "The joint distribution drifts. MI ESTIMATED FROM FINITE SAMPLES IS BIASED UPWARD, and the bias is worst at N ~ 10-30 - precisely where signal-value claims are least verifiable, which means the naive MI screen will nominate signals that carry no information at all. Mitigation: the Kraskov-Stogbauer-Grassberger estimator (asymptotically unbiased) rather than the small Miller-Madow correction, plus rolling re-estimation.",
"python_implementation": "dit; sklearn.feature_selection.mutual_info_classif (BIASED - use a KSG implementation for serious work)"
},
{
"name": "Kelly criterion / log-optimal growth",
"source_domain": "Information theory",
"citation": "Kelly (1956), Bell System Technical Journal 35(4): 917-926, doi:10.1002/j.1538-7305.1956.tb03809.x [T1]",
"transfer_mechanism": "Reference framework for sizing; identifies the maximum achievable exponential GROWTH RATE of capital with the mutual information between the bettor's private signal and the realized outcome. IMPORTANT EXCLUSION: one source states the growth-rate-optimal bet fraction is f* = I(X;Y)/H(X), attributing it to Cover & Thomas. This is a CATEGORY ERROR - Kelly's identity equates an achievable growth RATE (a quantity in bits or nats per bet) with mutual information, whereas the optimal bet FRACTION is a dimensionless capital share and a different object entirely. The formula is excluded from the merged report; the correct rate identity and the signal-RANKING use survive, the sizing rule does not.",
"transfer_risk": "Assumes a stationary distribution and that the true probability is known; MISESTIMATING P CAUSES CATASTROPHIC OVER-BETTING rather than mild inefficiency - it is the mechanism by which a positive-expectation strategy reaches zero. Separately and structurally, Kelly optimizes long-run geometric growth over many periods, which is NOT the objective of a fixed-multiple target under a hard deadline; both reports treat it as the indispensable starting reference and the wrong optimand for this specific problem.",
"python_implementation": "Manual (roughly ten readable lines of numpy); cvxpy for constrained sizing"
},
{
"name": "Multi-dimensional / log-optimal portfolio Kelly",
"source_domain": "Information theory",
"citation": "Cover & Thomas (2006), Elements of Information Theory, 2nd ed., Ch. 6, doi:10.1002/047174882X [T2]",
"transfer_mechanism": "Simultaneous allocation across several correlated bets.",
"transfer_risk": "Mathematically delicate; the source states it is rarely advisable without simulation. Correlated-signal structure is the binding difficulty.",
"python_implementation": "Numerical optimization; cvxpy"
},
{
"name": "KL divergence as a contract screen; entropy pooling",
"source_domain": "Information theory",
"citation": "Kullback & Leibler (1951), Annals of Mathematical Statistics 22: 79-86, doi:10.1214/aoms/1177729694 [T1]; entropy pooling in finance, Meucci (2010), 'Fully Flexible Views' [T4]",
"transfer_mechanism": "D_KL(p-hat || m) between subjective probability and market mid-price screens contracts for expected log-growth; entropy pooling projects a subjective view onto the market-implied prior under arbitrary constraints, generalizing Black-Litterman.",
"transfer_risk": "A TRAP THAT MUST BE STATED WITH THE TECHNIQUE: the growth identity holds ONLY WHEN p-hat IS THE TRUE PROBABILITY. Where p-hat merely DIFFERS from m, large divergence signals large expected LOSS exactly as readily as large expected gain. As written in the source, the screen licenses 'bet wherever you disagree with the market' - which is the precise failure the calibration machinery exists to prevent. USABLE ONLY AS A SECOND FILTER DOWNSTREAM OF DEMONSTRATED CALIBRATION, never as a standalone entry criterion. Entropy pooling additionally fails when view constraints are jointly infeasible.",
"python_implementation": "cvxpy projection; scipy.stats.entropy; ~50 lines for entropy pooling"
},
{
"name": "Maximum-entropy priors",
"source_domain": "Information theory",
"citation": "Jaynes (1957), Physical Review 106: 620-630, doi:10.1103/PhysRev.106.620 [T1]",
"transfer_mechanism": "Least-committal distribution consistent with known moments, for signal sources that are genuinely model-free.",
"transfer_risk": null,
"python_implementation": "scipy.optimize under moment constraints"
},
{
"name": "Fano's inequality",
"source_domain": "Information theory",
"citation": "Fano (1961), Transmission of Information, MIT Press [T2] - cited in one source's body and table with NO bibliographic record anywhere in that file",
"transfer_mechanism": "Lower-bounds misclassification probability given I(X;Y), hence a lower bound on the risk of the trader's bets.",
"transfer_risk": "Finite-sample MI bias propagates directly into the bound, making it optimistic.",
"python_implementation": "Manual computation"
},
{
"name": "Group sequential design and the Pocock boundary",
"source_domain": "Clinical-trial methodology",
"citation": "Pocock (1977), Biometrika 64(2): 191-199, doi:10.1093/biomet/64.2.191 [T1]",
"transfer_mechanism": "Pre-specified interim analyses with equal alpha spent per look, so that scheduled P&L reviews do not inflate the false-positive rate. This addresses the question no other domain in the inventory does: how often can a trader check P&L and still trust the verdict?",
"transfer_risk": "The number of looks must be fixed in advance and is itself a design parameter; a trader who peeks off-schedule invalidates the boundary. Mitigation: a self-binding SOFTWARE schedule rather than intent.",
"python_implementation": "gsDesign (R) port; manual boundary tables"
},
{
"name": "O'Brien-Fleming boundary",
"source_domain": "Clinical-trial methodology",
"citation": "O'Brien & Fleming (1979), Biometrics 35(3): 549-556, doi:10.2307/2530245 [T1]",
"transfer_mechanism": "A very conservative early boundary that spends almost no alpha at the first looks - matched to a setting where an early false positive is the expensive error.",
"transfer_risk": "Same fixed-look-count requirement as Pocock.",
"python_implementation": "gsDesign boundaries; scipy.stats.norm"
},
{
"name": "Lan-DeMets alpha-spending function",
"source_domain": "Clinical-trial methodology",
"citation": "Lan & DeMets (1983), Biometrika 70(3): 659-663, doi:10.1093/biomet/70.3.659 [T1]",
"transfer_mechanism": "Continuous alpha-spending that removes the requirement to fix the number of looks in advance; controls cumulative type-I error at 0.05 across all interim reviews. CONTEXT: one source computes 1 - 0.95^13 = 0.49 for K = 13 weekly checks; the arithmetic is right but the bound is for 13 INDEPENDENT tests, and interim looks at accumulating data are strongly positively correlated, so true inflation is materially lower. The direction survives - unstructured repeated peeking at P&L inflates false-positive rates substantially - but the 0.49 magnitude is [T6] and no replacement figure was invented.",
"transfer_risk": "Patient outcomes are independent; trade returns are serially autocorrelated, so the nominal boundary understates true spending. Same effective-sample-size remedy as SPRT.",
"python_implementation": "scipy.stats.norm spending function; gsDesign/rpact port"
},
{
"name": "Information-time versus calendar-time alpha spending",
"source_domain": "Clinical-trial methodology",
"citation": "Lan & DeMets (1989), Statistics in Medicine 8(10): 1191-1198, doi:10.1002/sim.4780081003 [T1]",
"transfer_mechanism": "THE MOST OPERATIONALLY CONSEQUENTIAL ITEM IN THIS DOMAIN AND ONE NO OTHER REPORT MAKES: with 90 days and probably <= 30 trades, calendar-time fraction (days elapsed / 90) and information-time fraction (effective sample size / target ESS) diverge sharply, because serial correlation makes the effective sample smaller than the trade count implies. Spend alpha on INFORMATION time. Practical effect: a trader 60 days into the experiment has spent far less than two-thirds of the available alpha.",
"transfer_risk": "The definition of information time is itself sensitive - it requires an effective-sample-size estimate that serial correlation makes uncertain.",
"python_implementation": "Manual; rpact"
},
{
"name": "Haybittle-Peto boundary",
"source_domain": "Clinical-trial methodology",
"citation": "Haybittle (1971); Peto et al. (1976) [T6] - the citing source flags its own appropriateness claim as author inference and both citations as requiring confirmation; neither has a complete bibliographic record",
"transfer_mechanism": "All interim looks evaluated at alpha ~ 0.001 with full alpha reserved for the final analysis - the most aggressive available protection against declaring edge early on noise, appropriate here because the cost of an early false positive (declaring the strategy works and increasing risk on the strength of noise) dominates the cost of a delayed decision.",
"transfer_risk": "Same fixed-look-count requirement as Pocock and O'Brien-Fleming.",
"python_implementation": "Manual boundary"
},
{
"name": "Pre-registration and protocol lock",
"source_domain": "Clinical-trial methodology",
"citation": "ClinicalTrials.gov guidance; FDA Modernization Act (1997) [T5]",
"transfer_mechanism": "A written protocol before launch binding three components: the parameterized hypothesis with all parameters bound, the a priori stopping rule as an alpha-spending function, and THE LOOK-ELSEWHERE CORRECTION recording how many candidate strategies were considered before this one. That third component is the one retail practice universally omits and the one that determines whether the final result means anything.",
"transfer_risk": "OVER-ENGINEERED AT THIS CAPITAL SCALE - institutional pre-registration cost is amortized across millions of dollars, and at USD 100 the human-attention cost is disproportionate. The resolution is precise: THE TRANSFER IS AT THE DISCIPLINE, NOT AT THE REGISTRY - a written protocol before launch, not a registry submission. A solo trader can also quietly revise the protocol; mitigation is a locked, timestamped document.",
"python_implementation": "Locked PDF; version control"
},
{
"name": "CONSORT reporting standard",
"source_domain": "Clinical-trial methodology",
"citation": "Schulz, Altman & Moher (2010), BMJ 340: c332, doi:10.1136/bmj.c332 [T1]",
"transfer_mechanism": "A reporting template ensuring the final write-up states what was pre-specified, what was changed and what was excluded.",
"transfer_risk": "Protocol flexibility under stress - the standard constrains reporting, not behaviour.",
"python_implementation": "Manual discipline"
},
{
"name": "Bayesian sequential design (posterior predictive)",
"source_domain": "Clinical-trial methodology",
"citation": "Spiegelhalter, Abrams & Myles (2004), Bayesian Approaches to Clinical Trials and Health-Care Evaluation [T2]; Jennison & Turnbull (2000) [T2]",
"transfer_mechanism": "Declare success when P(edge > 0 | data) > 0.95 under a beta-binomial conjugate posterior. THE OPERATIONALLY CLEANER ALTERNATIVE to alpha-spending for a single trader: the conjugate posterior is a two-line computation, requires no boundary tables, and handles unscheduled looks natively. Stopping-rule asymmetry should be preserved from the clinical analogue - suspend on a low P&L threshold at a 30-trade sample, but do not CONFIRM edge on a high threshold until the final analysis.",
"transfer_risk": "Model misspecification - the conjugate posterior assumes a fixed win probability across a sample where the underlying rate may be drifting. Mitigation: posterior predictive checks. THE META-RISK THIS DOMAIN NAMES AND NO OTHER DOES: high methodological sophistication creates the ILLUSION of robust edge, because every score is good, every credibility factor is high and every interim peek passes. The remedy is identical in both domains - at conclusion, perform one final fully-specified test, and accept that the answer can be 'no edge' without the prior work having been wasted.",
"python_implementation": "scipy.stats.beta; PyMC. NOTE: the symbol 'betaind from scipy' cited by one source DOES NOT EXIST; the correct symbol is scipy.stats.beta, which that source's own table uses elsewhere."
}
],
"python_libraries": [
{
"package": "numpy",
"version": "2.5.1",
"latest_release_date": "2026-07-04",
"license_spdx": "BSD-3-Clause",
"repo_url": "https://github.com/numpy/numpy",
"stars": 32469,
"open_issues": 2317,
"maintenance_status": "active",
"architecture_layer": "Substrate",
"verified_via": "unverified - neither surviving source captured a single PyPI or GitHub API response (no JSON excerpts, no HTTP status codes, no per-record retrieval timestamps, no ETags), so every version string, release date, star count and open-issue count below is an unsourced assertion requiring independent re-verification via `pip index versions <name>` or https://pypi.org/pypi/<name>/json. The package SELECTION and architecture-layer assignments survive the audit; the metadata precision does not. Claims of `verified_via: PyPI API` in a source appendix are not reproduced here."
},
{
"package": "scipy",
"version": "1.18.0",
"latest_release_date": "2026-06-19",
"license_spdx": "BSD-3-Clause",
"repo_url": "https://github.com/scipy/scipy",
"stars": 14875,
"open_issues": 1846,
"maintenance_status": "active",
"architecture_layer": "Substrate",
"verified_via": "unverified - neither surviving source captured a single PyPI or GitHub API response (no JSON excerpts, no HTTP status codes, no per-record retrieval timestamps, no ETags), so every version string, release date, star count and open-issue count below is an unsourced assertion requiring independent re-verification via `pip index versions <name>` or https://pypi.org/pypi/<name>/json. The package SELECTION and architecture-layer assignments survive the audit; the metadata precision does not. Claims of `verified_via: PyPI API` in a source appendix are not reproduced here."
},
{
"package": "pandas",
"version": "3.0.5",
"latest_release_date": "2026-07-22",
"license_spdx": "BSD-3-Clause",
"repo_url": "https://github.com/pandas-dev/pandas",
"stars": 49388,
"open_issues": 2924,
"maintenance_status": "active",
"architecture_layer": "Substrate",
"verified_via": "unverified - neither surviving source captured a single PyPI or GitHub API response (no JSON excerpts, no HTTP status codes, no per-record retrieval timestamps, no ETags), so every version string, release date, star count and open-issue count below is an unsourced assertion requiring independent re-verification via `pip index versions <name>` or https://pypi.org/pypi/<name>/json. The package SELECTION and architecture-layer assignments survive the audit; the metadata precision does not. Claims of `verified_via: PyPI API` in a source appendix are not reproduced here."
},
{
"package": "polars",
"version": "1.43.2",
"latest_release_date": "2026-08-01",
"license_spdx": "MIT",
"repo_url": "https://github.com/pola-rs/polars",
"stars": 39156,
"open_issues": 2846,
"maintenance_status": "active",
"architecture_layer": "Substrate / feature computation",
"verified_via": "unverified - neither surviving source captured a single PyPI or GitHub API response (no JSON excerpts, no HTTP status codes, no per-record retrieval timestamps, no ETags), so every version string, release date, star count and open-issue count below is an unsourced assertion requiring independent re-verification via `pip index versions <name>` or https://pypi.org/pypi/<name>/json. The package SELECTION and architecture-layer assignments survive the audit; the metadata precision does not. Claims of `verified_via: PyPI API` in a source appendix are not reproduced here. | ROW NOTE: Version conflict 1.43.2 vs 1.21.0. Resolved to the higher, but flagged: the asserted release date is the same calendar day as the research date, and a zero-day-old release is exactly the shape of a value generated to satisfy a recency rule rather than observed."
},
{
"package": "pyarrow",
"version": "25.0.0",
"latest_release_date": "2026-07-10",
"license_spdx": "Apache-2.0",
"repo_url": "https://github.com/apache/arrow",
"stars": 16969,
"open_issues": 2554,
"maintenance_status": "active",
"architecture_layer": "Substrate / columnar interchange",
"verified_via": "unverified - neither surviving source captured a single PyPI or GitHub API response (no JSON excerpts, no HTTP status codes, no per-record retrieval timestamps, no ETags), so every version string, release date, star count and open-issue count below is an unsourced assertion requiring independent re-verification via `pip index versions <name>` or https://pypi.org/pypi/<name>/json. The package SELECTION and architecture-layer assignments survive the audit; the metadata precision does not. Claims of `verified_via: PyPI API` in a source appendix are not reproduced here."
},
{
"package": "duckdb",
"version": "1.5.5",
"latest_release_date": "2026-07-22",
"license_spdx": "MIT",
"repo_url": "https://github.com/duckdb/duckdb",
"stars": 24500,
"open_issues": 420,
"maintenance_status": "active",
"architecture_layer": "Storage",
"verified_via": "unverified - neither surviving source captured a single PyPI or GitHub API response (no JSON excerpts, no HTTP status codes, no per-record retrieval timestamps, no ETags), so every version string, release date, star count and open-issue count below is an unsourced assertion requiring independent re-verification via `pip index versions <name>` or https://pypi.org/pypi/<name>/json. The package SELECTION and architecture-layer assignments survive the audit; the metadata precision does not. Claims of `verified_via: PyPI API` in a source appendix are not reproduced here. | ROW NOTE: Star count resolved to 24,500 against the main project repo; the competing figure of 174 is correct for the duckdb-python client sub-repo but misleads a reader scanning a maintenance-health column."
},
{
"package": "yfinance",
"version": "1.5.2",
"latest_release_date": "2026-07-23",
"license_spdx": "Apache-2.0",
"repo_url": "https://github.com/ranaroussi/yfinance",
"stars": 24856,
"open_issues": 169,
"maintenance_status": "active",
"architecture_layer": "Ingestion - Yahoo EOD/OHLCV",
"verified_via": "unverified - neither surviving source captured a single PyPI or GitHub API response (no JSON excerpts, no HTTP status codes, no per-record retrieval timestamps, no ETags), so every version string, release date, star count and open-issue count below is an unsourced assertion requiring independent re-verification via `pip index versions <name>` or https://pypi.org/pypi/<name>/json. The package SELECTION and architecture-layer assignments survive the audit; the metadata precision does not. Claims of `verified_via: PyPI API` in a source appendix are not reproduced here. | ROW NOTE: Its constraint is legal, not technical: Yahoo's terms license personal, non-commercial use only. The single most-used ingestion package in retail quant is the one operating furthest outside its provider's terms, and it is the only data source in either audit flagged point_in_time=false, survivorship_bias_free=false AND tos_restricts_automation=true."
},
{
"package": "pandas-datareader",
"version": "0.11.1",
"latest_release_date": "2026-06-24",
"license_spdx": "BSD-3-Clause",
"repo_url": "https://github.com/pydata/pandas-datareader",
"stars": 3226,
"open_issues": 145,
"maintenance_status": "slowing",
"architecture_layer": "Ingestion - FRED / World Bank / OECD",
"verified_via": "unverified - neither surviving source captured a single PyPI or GitHub API response (no JSON excerpts, no HTTP status codes, no per-record retrieval timestamps, no ETags), so every version string, release date, star count and open-issue count below is an unsourced assertion requiring independent re-verification via `pip index versions <name>` or https://pypi.org/pypi/<name>/json. The package SELECTION and architecture-layer assignments survive the audit; the metadata precision does not. Claims of `verified_via: PyPI API` in a source appendix are not reproduced here. | ROW NOTE: Marked active on a 2026-06-24 release while the same source describes the project as having 'the first release in over a year' - recorded here as slowing. Its Yahoo path broke in 2020 and has not returned. Fallback is trivial: vendor the FRED REST calls through httpx directly."
},
{
"package": "ccxt",
"version": "4.5.70",
"latest_release_date": "2026-07-29",
"license_spdx": "MIT",
"repo_url": "https://github.com/ccxt/ccxt",
"stars": 43470,
"open_issues": 937,
"maintenance_status": "active",
"architecture_layer": "Ingestion - crypto CEX/DEX",
"verified_via": "unverified - neither surviving source captured a single PyPI or GitHub API response (no JSON excerpts, no HTTP status codes, no per-record retrieval timestamps, no ETags), so every version string, release date, star count and open-issue count below is an unsourced assertion requiring independent re-verification via `pip index versions <name>` or https://pypi.org/pypi/<name>/json. The package SELECTION and architecture-layer assignments survive the audit; the metadata precision does not. Claims of `verified_via: PyPI API` in a source appendix are not reproduced here. | ROW NOTE: Unifies 105+ crypto venues behind one API. Commercial boundary: CCXT Pro (WebSocket streaming) is a paid product; the open-source package is REST-only."
},
{
"package": "ib-async",
"version": "2.1.0",
"latest_release_date": "2025-12-08",
"license_spdx": "BSD-2-Clause",
"repo_url": "https://github.com/ib-api-reloaded/ib_async",
"stars": 1707,
"open_issues": 89,
"maintenance_status": "active",
"architecture_layer": "Ingestion + live execution - IBKR / ForecastEx",
"verified_via": "unverified - neither surviving source captured a single PyPI or GitHub API response (no JSON excerpts, no HTTP status codes, no per-record retrieval timestamps, no ETags), so every version string, release date, star count and open-issue count below is an unsourced assertion requiring independent re-verification via `pip index versions <name>` or https://pypi.org/pypi/<name>/json. The package SELECTION and architecture-layer assignments survive the audit; the metadata precision does not. Claims of `verified_via: PyPI API` in a source appendix are not reproduced here. | ROW NOTE: The only actively maintained Python framework speaking IBKR's native protocol. Repository conflict resolved on mechanism: the package was renamed and the repository moved after the original maintainer's death in early 2024, superseding the abandoned ib_insync. The competing erdewit/ib_async URL is the pre-transfer location. License is unsourced by either report."
},
{
"package": "polygon-api-client",
"version": "1.16.3",
"latest_release_date": "2025-10-30",
"license_spdx": "MIT",
"repo_url": "https://github.com/polygon-io/client-python",
"stars": 1490,
"open_issues": 19,
"maintenance_status": "active",
"architecture_layer": "Ingestion - US equities / options chains",
"verified_via": "unverified - neither surviving source captured a single PyPI or GitHub API response (no JSON excerpts, no HTTP status codes, no per-record retrieval timestamps, no ETags), so every version string, release date, star count and open-issue count below is an unsourced assertion requiring independent re-verification via `pip index versions <name>` or https://pypi.org/pypi/<name>/json. The package SELECTION and architecture-layer assignments survive the audit; the metadata precision does not. Claims of `verified_via: PyPI API` in a source appendix are not reproduced here. | ROW NOTE: Retained in preference to a claimed successor package `massive` 2.8.0, whose entire existence rests on a README-sourced rebrand claim with an asserted rebrand date identical to this package's asserted release date."
},
{
"package": "statsmodels",
"version": "0.14.6",
"latest_release_date": "2025-12-05",
"license_spdx": "BSD-3-Clause",
"repo_url": "https://github.com/statsmodels/statsmodels",
"stars": 11546,
"open_issues": 2888,
"maintenance_status": "active",
"architecture_layer": "Econometrics + state space",
"verified_via": "unverified - neither surviving source captured a single PyPI or GitHub API response (no JSON excerpts, no HTTP status codes, no per-record retrieval timestamps, no ETags), so every version string, release date, star count and open-issue count below is an unsourced assertion requiring independent re-verification via `pip index versions <name>` or https://pypi.org/pypi/<name>/json. The package SELECTION and architecture-layer assignments survive the audit; the metadata precision does not. Claims of `verified_via: PyPI API` in a source appendix are not reproduced here. | ROW NOTE: Version agreed across both sources; release date conflicted (2025-12-05 vs 2026-04-10) and the same version cannot have two release dates. Its statespace submodule covers Kalman filtering and structural time series with tighter integration than any standalone filter package, which is why filterpy is excluded and pykalman is optional."
},
{
"package": "arch",
"version": "8.0.0",
"latest_release_date": "2025-10-21",
"license_spdx": "NCSA",
"repo_url": "https://github.com/bashtage/arch",
"stars": 1548,
"open_issues": 51,
"maintenance_status": "active",
"architecture_layer": "GARCH / volatility / unit root",
"verified_via": "unverified - neither surviving source captured a single PyPI or GitHub API response (no JSON excerpts, no HTTP status codes, no per-record retrieval timestamps, no ETags), so every version string, release date, star count and open-issue count below is an unsourced assertion requiring independent re-verification via `pip index versions <name>` or https://pypi.org/pypi/<name>/json. The package SELECTION and architecture-layer assignments survive the audit; the metadata precision does not. Claims of `verified_via: PyPI API` in a source appendix are not reproduced here. | ROW NOTE: Version (8.0.0 vs 7.2.0) and license (NCSA vs MIT) both conflicted. Resolved on the explicit pyproject.toml license declaration; NCSA is permissive and functionally MIT-equivalent for this use, which explains the approximation. IMPORTANT: arch.bootstrap has provided first-class SPA, StepM and MCS classes for years - Hansen's SPA test is directly callable and should NOT be reimplemented, contrary to one source's claim."
},
{
"package": "linearmodels",
"version": "7.0",
"latest_release_date": "2025-10-21",
"license_spdx": "NCSA",
"repo_url": "https://github.com/bashtage/linearmodels",
"stars": 1060,
"open_issues": 56,
"maintenance_status": "active",
"architecture_layer": "Panel / IV / asset pricing",
"verified_via": "unverified - neither surviving source captured a single PyPI or GitHub API response (no JSON excerpts, no HTTP status codes, no per-record retrieval timestamps, no ETags), so every version string, release date, star count and open-issue count below is an unsourced assertion requiring independent re-verification via `pip index versions <name>` or https://pypi.org/pypi/<name>/json. The package SELECTION and architecture-layer assignments survive the audit; the metadata precision does not. Claims of `verified_via: PyPI API` in a source appendix are not reproduced here. | ROW NOTE: Same maintainer and same license-declaration mechanism as arch; the MIT-vs-NCSA conflict resolves identically."
},
{
"package": "pmdarima",
"version": "2.1.1",
"latest_release_date": "2025-11-17",
"license_spdx": "MIT",
"repo_url": "https://github.com/alkaline-ml/pmdarima",
"stars": 1732,
"open_issues": 64,
"maintenance_status": "slowing",
"architecture_layer": "Auto-ARIMA (reference only)",
"verified_via": "unverified - neither surviving source captured a single PyPI or GitHub API response (no JSON excerpts, no HTTP status codes, no per-record retrieval timestamps, no ETags), so every version string, release date, star count and open-issue count below is an unsourced assertion requiring independent re-verification via `pip index versions <name>` or https://pypi.org/pypi/<name>/json. The package SELECTION and architecture-layer assignments survive the audit; the metadata precision does not. Claims of `verified_via: PyPI API` in a source appendix are not reproduced here. | ROW NOTE: One release in the trailing 365 days, with its lead maintainer having redirected primary effort to statsforecast. Retained as a reference implementation; for new code use statsforecast.AutoARIMA."
},
{
"package": "pykalman",
"version": "0.11.2",
"latest_release_date": "2026-01-31",
"license_spdx": "BSD-3-Clause",
"repo_url": "https://github.com/pykalman/pykalman",
"stars": 1327,
"open_issues": 85,
"maintenance_status": "slowing",
"architecture_layer": "State space (optional)",
"verified_via": "unverified - neither surviving source captured a single PyPI or GitHub API response (no JSON excerpts, no HTTP status codes, no per-record retrieval timestamps, no ETags), so every version string, release date, star count and open-issue count below is an unsourced assertion requiring independent re-verification via `pip index versions <name>` or https://pypi.org/pypi/<name>/json. The package SELECTION and architecture-layer assignments survive the audit; the metadata precision does not. Claims of `verified_via: PyPI API` in a source appendix are not reproduced here. | ROW NOTE: Survives recency but earns no place in the reference architecture; statsmodels.tsa.statespace covers the same ground with better integration."
},
{
"package": "sktime",
"version": "1.1.0",
"latest_release_date": "2026-07-28",
"license_spdx": "BSD-3-Clause",
"repo_url": "https://github.com/sktime/sktime",
"stars": 9896,
"open_issues": 2371,
"maintenance_status": "active",
"architecture_layer": "Forecasting framework",
"verified_via": "unverified - neither surviving source captured a single PyPI or GitHub API response (no JSON excerpts, no HTTP status codes, no per-record retrieval timestamps, no ETags), so every version string, release date, star count and open-issue count below is an unsourced assertion requiring independent re-verification via `pip index versions <name>` or https://pypi.org/pypi/<name>/json. The package SELECTION and architecture-layer assignments survive the audit; the metadata precision does not. Claims of `verified_via: PyPI API` in a source appendix are not reproduced here. | ROW NOTE: Highest-velocity framework in the survey: a unified fit/predict API over 500+ models spanning classical, ML, deep learning and foundation models. PICK ONE OF sktime OR darts as the primary API - running both against the same problem doubles the surface area for no gain."
},
{
"package": "statsforecast",
"version": "2.1.1",
"latest_release_date": "2026-07-16",
"license_spdx": "Apache-2.0",
"repo_url": "https://github.com/Nixtla/statsforecast",
"stars": 4854,
"open_issues": 139,
"maintenance_status": "active",
"architecture_layer": "Forecasting - fast statistical",
"verified_via": "unverified - neither surviving source captured a single PyPI or GitHub API response (no JSON excerpts, no HTTP status codes, no per-record retrieval timestamps, no ETags), so every version string, release date, star count and open-issue count below is an unsourced assertion requiring independent re-verification via `pip index versions <name>` or https://pypi.org/pypi/<name>/json. The package SELECTION and architecture-layer assignments survive the audit; the metadata precision does not. Claims of `verified_via: PyPI API` in a source appendix are not reproduced here. | ROW NOTE: Version conflict 2.1.1 vs 2.0.3 resolved to the higher. The right default at this scale: lighter than darts, faster than statsmodels on the same models. Its claimed 20x speedup over pmdarima is a vendor benchmark."
},
{
"package": "darts",
"version": "0.46.1",
"latest_release_date": "2026-07-20",
"license_spdx": "Apache-2.0",
"repo_url": "https://github.com/unit8co/darts",
"stars": 9480,
"open_issues": 215,
"maintenance_status": "active",
"architecture_layer": "Forecasting - unified + neural",
"verified_via": "unverified - neither surviving source captured a single PyPI or GitHub API response (no JSON excerpts, no HTTP status codes, no per-record retrieval timestamps, no ETags), so every version string, release date, star count and open-issue count below is an unsourced assertion requiring independent re-verification via `pip index versions <name>` or https://pypi.org/pypi/<name>/json. The package SELECTION and architecture-layer assignments survive the audit; the metadata precision does not. Claims of `verified_via: PyPI API` in a source appendix are not reproduced here. | ROW NOTE: Version conflict 0.46.1 vs 0.31.0; a 15-minor-version gap is implausible as noise and the higher figure is carried. Comparable in scope to sktime with stronger neural and probabilistic support, at the cost of a PyTorch-Lightning dependency and a large install."
},
{
"package": "gluonts",
"version": "0.17.0",
"latest_release_date": "2026-07-31",
"license_spdx": "Apache-2.0",
"repo_url": "https://github.com/awslabs/gluonts",
"stars": 5221,
"open_issues": 470,
"maintenance_status": "slowing",
"architecture_layer": "Forecasting - probabilistic (DeepAR)",
"verified_via": "unverified - neither surviving source captured a single PyPI or GitHub API response (no JSON excerpts, no HTTP status codes, no per-record retrieval timestamps, no ETags), so every version string, release date, star count and open-issue count below is an unsourced assertion requiring independent re-verification via `pip index versions <name>` or https://pypi.org/pypi/<name>/json. The package SELECTION and architecture-layer assignments survive the audit; the metadata precision does not. Claims of `verified_via: PyPI API` in a source appendix are not reproduced here. | ROW NOTE: Reference implementation of DeepAR; cadence is slowing. Reach for darts first."
},
{
"package": "neuralforecast",
"version": "3.2.0",
"latest_release_date": "2026-07-10",
"license_spdx": "Apache-2.0",
"repo_url": "https://github.com/Nixtla/neuralforecast",
"stars": null,
"open_issues": null,
"maintenance_status": "active",
"architecture_layer": "Forecasting - deep learning",
"verified_via": "unverified - neither surviving source captured a single PyPI or GitHub API response (no JSON excerpts, no HTTP status codes, no per-record retrieval timestamps, no ETags), so every version string, release date, star count and open-issue count below is an unsourced assertion requiring independent re-verification via `pip index versions <name>` or https://pypi.org/pypi/<name>/json. The package SELECTION and architecture-layer assignments survive the audit; the metadata precision does not. Claims of `verified_via: PyPI API` in a source appendix are not reproduced here. | ROW NOTE: Star and issue counts are n/a in the source and recorded as null rather than guessed. Not needed for the doubling experiment; included for forward extension."
},
{
"package": "hierarchicalforecast",
"version": "1.5.1",
"latest_release_date": "2026-03-04",
"license_spdx": "Apache-2.0",
"repo_url": "https://github.com/Nixtla/hierarchicalforecast",
"stars": 752,
"open_issues": 7,
"maintenance_status": "active",
"architecture_layer": "Forecasting - hierarchical reconciliation",
"verified_via": "unverified - neither surviving source captured a single PyPI or GitHub API response (no JSON excerpts, no HTTP status codes, no per-record retrieval timestamps, no ETags), so every version string, release date, star count and open-issue count below is an unsourced assertion requiring independent re-verification via `pip index versions <name>` or https://pypi.org/pypi/<name>/json. The package SELECTION and architecture-layer assignments survive the audit; the metadata precision does not. Claims of `verified_via: PyPI API` in a source appendix are not reproduced here. | ROW NOTE: THE ONLY LIBRARY IN THE SURVEYED UNIVERSE THAT ADDRESSES HIERARCHICAL RECONCILIATION (BottomUp, TopDown, MinTrace, ERM, PERMBU, conformal methods). Single-source, retained because it answers a required capability nothing else covers."
},
{
"package": "pymc",
"version": "6.2.0",
"latest_release_date": "2026-07-23",
"license_spdx": "Apache-2.0",
"repo_url": "https://github.com/pymc-devs/pymc",
"stars": 9695,
"open_issues": 479,
"maintenance_status": "active",
"architecture_layer": "Bayesian inference (primary PPL)",
"verified_via": "unverified - neither surviving source captured a single PyPI or GitHub API response (no JSON excerpts, no HTTP status codes, no per-record retrieval timestamps, no ETags), so every version string, release date, star count and open-issue count below is an unsourced assertion requiring independent re-verification via `pip index versions <name>` or https://pypi.org/pypi/<name>/json. The package SELECTION and architecture-layer assignments survive the audit; the metadata precision does not. Claims of `verified_via: PyPI API` in a source appendix are not reproduced here. | ROW NOTE: GENUINELY CONTESTED: 6.2.0 versus a contemporaneous 5.17.0 claim. The higher is carried because its source supplies a mechanism (6.x as a stabilized major API rewrite) rather than a bare number, but a full major-version divergence between contemporaneous sources must be resolved at install time. If 6.x is real, expect breaking API changes against every PyMC tutorial written before it. PICK ONE PRIMARY PPL - running two is a maintenance tax with no analytical payoff at this scale."
},
{
"package": "pytensor",
"version": "3.2.3",
"latest_release_date": "2026-07-25",
"license_spdx": "BSD-3-Clause",
"repo_url": "https://github.com/pymc-devs/pytensor",
"stars": null,
"open_issues": null,
"maintenance_status": "active",
"architecture_layer": "Bayesian - symbolic compiler",
"verified_via": "unverified - neither surviving source captured a single PyPI or GitHub API response (no JSON excerpts, no HTTP status codes, no per-record retrieval timestamps, no ETags), so every version string, release date, star count and open-issue count below is an unsourced assertion requiring independent re-verification via `pip index versions <name>` or https://pypi.org/pypi/<name>/json. The package SELECTION and architecture-layer assignments survive the audit; the metadata precision does not. Claims of `verified_via: PyPI API` in a source appendix are not reproduced here. | ROW NOTE: Star and issue counts n/a in the source; null rather than guessed. The Theano/Aesara successor underlying PyMC."
},
{
"package": "arviz",
"version": "1.2.0",
"latest_release_date": "2026-06-12",
"license_spdx": "Apache-2.0",
"repo_url": "https://github.com/arviz-devs/arviz",
"stars": null,
"open_issues": null,
"maintenance_status": "active",
"architecture_layer": "Bayesian - posterior diagnostics",
"verified_via": "unverified - neither surviving source captured a single PyPI or GitHub API response (no JSON excerpts, no HTTP status codes, no per-record retrieval timestamps, no ETags), so every version string, release date, star count and open-issue count below is an unsourced assertion requiring independent re-verification via `pip index versions <name>` or https://pypi.org/pypi/<name>/json. The package SELECTION and architecture-layer assignments survive the audit; the metadata precision does not. Claims of `verified_via: PyPI API` in a source appendix are not reproduced here. | ROW NOTE: Star and issue counts n/a in the source; null rather than guessed. Not a PPL - the diagnostics layer for one (ESS, R-hat, LOO, WAIC, trace plots, posterior predictive checks). Install it alongside whichever sampler you choose, ALWAYS: a Bayesian forecast published without R-hat and ESS is an unaudited number."
},
{
"package": "numpyro",
"version": "0.21.0",
"latest_release_date": "2026-05-02",
"license_spdx": "Apache-2.0",
"repo_url": "https://github.com/pyro-ppl/numpyro",
"stars": 2730,
"open_issues": 68,
"maintenance_status": "active",
"architecture_layer": "Bayesian - JAX / GPU PPL",
"verified_via": "unverified - neither surviving source captured a single PyPI or GitHub API response (no JSON excerpts, no HTTP status codes, no per-record retrieval timestamps, no ETags), so every version string, release date, star count and open-issue count below is an unsourced assertion requiring independent re-verification via `pip index versions <name>` or https://pypi.org/pypi/<name>/json. The package SELECTION and architecture-layer assignments survive the audit; the metadata precision does not. Claims of `verified_via: PyPI API` in a source appendix are not reproduced here. | ROW NOTE: Two-source version agreement - among the strongest corroboration in the cluster. Take it if JAX is already resident on a GPU. Not necessary for a USD 100 experiment."
},
{
"package": "cmdstanpy",
"version": "1.3.0",
"latest_release_date": "2025-10-20",
"license_spdx": "BSD-3-Clause",
"repo_url": "https://github.com/stan-dev/cmdstanpy",
"stars": 198,
"open_issues": 29,
"maintenance_status": "active",
"architecture_layer": "Bayesian - Stan HMC/NUTS",
"verified_via": "unverified - neither surviving source captured a single PyPI or GitHub API response (no JSON excerpts, no HTTP status codes, no per-record retrieval timestamps, no ETags), so every version string, release date, star count and open-issue count below is an unsourced assertion requiring independent re-verification via `pip index versions <name>` or https://pypi.org/pypi/<name>/json. The package SELECTION and architecture-layer assignments survive the audit; the metadata precision does not. Claims of `verified_via: PyPI API` in a source appendix are not reproduced here. | ROW NOTE: Two-source version agreement. Take it if you want Stan's HMC/NUTS implementation and its documentation, which remains the best in the field."
},
{
"package": "vectorbt",
"version": "1.1.0",
"latest_release_date": "2026-07-05",
"license_spdx": "Apache-2.0 AND Commons-Clause",
"repo_url": "https://github.com/polakowo/vectorbt",
"stars": 8515,
"open_issues": 136,
"maintenance_status": "active",
"architecture_layer": "Backtesting - vectorized sweeps",
"verified_via": "unverified - neither surviving source captured a single PyPI or GitHub API response (no JSON excerpts, no HTTP status codes, no per-record retrieval timestamps, no ETags), so every version string, release date, star count and open-issue count below is an unsourced assertion requiring independent re-verification via `pip index versions <name>` or https://pypi.org/pypi/<name>/json. The package SELECTION and architecture-layer assignments survive the audit; the metadata precision does not. Claims of `verified_via: PyPI API` in a source appendix are not reproduced here. | ROW NOTE: LICENSE CONFLICT RESOLVED TO THE MORE RESTRICTIVE: Apache-2.0 WITH the Commons Clause addendum, not bare Apache-2.0. The specific falsifiable claim beats the SPDX field GitHub returns, which does not represent addenda. Practical effect: free for individuals and organizations, but you may not sell a product or service whose value derives primarily from the software. Irrelevant for a personal experiment; material if the stack is ever packaged."
},
{
"package": "backtesting",
"version": "0.6.6",
"latest_release_date": "2026-07-22",
"license_spdx": "AGPL-3.0-or-later",
"repo_url": "https://github.com/kernc/backtesting.py",
"stars": 8745,
"open_issues": 61,
"maintenance_status": "active",
"architecture_layer": "Backtesting - event-driven research",
"verified_via": "unverified - neither surviving source captured a single PyPI or GitHub API response (no JSON excerpts, no HTTP status codes, no per-record retrieval timestamps, no ETags), so every version string, release date, star count and open-issue count below is an unsourced assertion requiring independent re-verification via `pip index versions <name>` or https://pypi.org/pypi/<name>/json. The package SELECTION and architecture-layer assignments survive the audit; the metadata precision does not. Claims of `verified_via: PyPI API` in a source appendix are not reproduced here. | ROW NOTE: GENUINE COPYLEFT: distributing a derivative, INCLUDING OVER A NETWORK, triggers source-disclosure obligations. Non-issue privately; disqualifying for a hosted service. Single-author maintenance."
},
{
"package": "bt",
"version": "1.2.0",
"latest_release_date": "2026-04-25",
"license_spdx": "MIT",
"repo_url": "https://github.com/pmorissette/bt",
"stars": 2954,
"open_issues": 83,
"maintenance_status": "active",
"architecture_layer": "Backtesting - tree/portfolio composition",
"verified_via": "unverified - neither surviving source captured a single PyPI or GitHub API response (no JSON excerpts, no HTTP status codes, no per-record retrieval timestamps, no ETags), so every version string, release date, star count and open-issue count below is an unsourced assertion requiring independent re-verification via `pip index versions <name>` or https://pypi.org/pypi/<name>/json. The package SELECTION and architecture-layer assignments survive the audit; the metadata precision does not. Claims of `verified_via: PyPI API` in a source appendix are not reproduced here. | ROW NOTE: Narrow but clean if you are composing weighted sleeves rather than trading signals."
},
{
"package": "nautilus-trader",
"version": "1.230.0",
"latest_release_date": "2026-06-29",
"license_spdx": "LGPL-3.0-or-later",
"repo_url": "https://github.com/nautechsystems/nautilus_trader",
"stars": 25180,
"open_issues": 82,
"maintenance_status": "active",
"architecture_layer": "Execution decision - event-driven, Rust core",
"verified_via": "unverified - neither surviving source captured a single PyPI or GitHub API response (no JSON excerpts, no HTTP status codes, no per-record retrieval timestamps, no ETags), so every version string, release date, star count and open-issue count below is an unsourced assertion requiring independent re-verification via `pip index versions <name>` or https://pypi.org/pypi/<name>/json. The package SELECTION and architecture-layer assignments survive the audit; the metadata precision does not. Claims of `verified_via: PyPI API` in a source appendix are not reproduced here. | ROW NOTE: Version conflict resolved to 1.230.0: the competing claim pairs a LOWER version (1.218.0) with a LATER date (2026-07-28), which is internally inconsistent. The most production-grade option and the steepest learning curve in the survey. It is the only engine here that runs the same code in backtest and live, which is why it belongs at the decision boundary - one source assigns it to execution and the other to backtesting, and both are right. Optional at USD 100 scale."
},
{
"package": "cvxpy",
"version": "1.9.2",
"latest_release_date": "2026-06-22",
"license_spdx": "Apache-2.0",
"repo_url": "https://github.com/cvxpy/cvxpy",
"stars": 6293,
"open_issues": 192,
"maintenance_status": "active",
"architecture_layer": "Sizing - convex programming",
"verified_via": "unverified - neither surviving source captured a single PyPI or GitHub API response (no JSON excerpts, no HTTP status codes, no per-record retrieval timestamps, no ETags), so every version string, release date, star count and open-issue count below is an unsourced assertion requiring independent re-verification via `pip index versions <name>` or https://pypi.org/pypi/<name>/json. The package SELECTION and architecture-layer assignments survive the audit; the metadata precision does not. Claims of `verified_via: PyPI API` in a source appendix are not reproduced here. | ROW NOTE: Two-source version agreement. The DSL underneath BOTH PyPortfolioOpt and Riskfolio-Lib - install it directly, because the moment you need a custom objective or constraint you are writing cvxpy anyway."
},
{
"package": "PyPortfolioOpt",
"version": "1.6.0",
"latest_release_date": "2026-02-26",
"license_spdx": "MIT",
"repo_url": "https://github.com/robertmartin8/PyPortfolioOpt",
"stars": 5922,
"open_issues": 109,
"maintenance_status": "slowing",
"architecture_layer": "Sizing - mean-variance / Black-Litterman",
"verified_via": "unverified - neither surviving source captured a single PyPI or GitHub API response (no JSON excerpts, no HTTP status codes, no per-record retrieval timestamps, no ETags), so every version string, release date, star count and open-issue count below is an unsourced assertion requiring independent re-verification via `pip index versions <name>` or https://pypi.org/pypi/<name>/json. The package SELECTION and architecture-layer assignments survive the audit; the metadata precision does not. Claims of `verified_via: PyPI API` in a source appendix are not reproduced here. | ROW NOTE: The most widely cited mean-variance and Black-Litterman implementation, at a slowing cadence. Keep it only if you want its specific Black-Litterman API."
},
{
"package": "Riskfolio-Lib",
"version": "7.3.0",
"latest_release_date": "2026-05-31",
"license_spdx": "BSD-3-Clause",
"repo_url": "https://github.com/dcajasn/Riskfolio-Lib",
"stars": 4420,
"open_issues": 28,
"maintenance_status": "active",
"architecture_layer": "Sizing - advanced risk measures",
"verified_via": "unverified - neither surviving source captured a single PyPI or GitHub API response (no JSON excerpts, no HTTP status codes, no per-record retrieval timestamps, no ETags), so every version string, release date, star count and open-issue count below is an unsourced assertion requiring independent re-verification via `pip index versions <name>` or https://pypi.org/pypi/<name>/json. The package SELECTION and architecture-layer assignments survive the audit; the metadata precision does not. Claims of `verified_via: PyPI API` in a source appendix are not reproduced here. | ROW NOTE: Two-source version agreement. The most capable single library in the survey: 26 convex risk measures, risk-parity variants, hierarchical clustering, nested clustered optimization, Worst-Case Mean-Variance, OWA, MVSK, Black-Litterman, entropy pooling, cardinality constraints. WARNING: a fully-qualified Kelly path of the form Riskfolio-Lib.optimization.mean_risk.portfolio_kelly was asserted by one source with zero documentation links and no version pin, and its own digest flags it as the least-sourced actionable line in the cluster - DO NOT TRUST THAT PATH."
},
{
"package": "skfolio",
"version": "0.20.1",
"latest_release_date": "2026-04-21",
"license_spdx": "BSD-3-Clause",
"repo_url": "https://github.com/skfolio/skfolio",
"stars": 2085,
"open_issues": 21,
"maintenance_status": "active",
"architecture_layer": "Sizing + CombinatorialPurgedCV / WalkForward",
"verified_via": "unverified - neither surviving source captured a single PyPI or GitHub API response (no JSON excerpts, no HTTP status codes, no per-record retrieval timestamps, no ETags), so every version string, release date, star count and open-issue count below is an unsourced assertion requiring independent re-verification via `pip index versions <name>` or https://pypi.org/pypi/<name>/json. The package SELECTION and architecture-layer assignments survive the audit; the metadata precision does not. Claims of `verified_via: PyPI API` in a source appendix are not reproduced here. | ROW NOTE: THE INTEGRATION POINT, and the single most useful finding in the Python cluster despite being single-source: the only portfolio library following the scikit-learn fit/predict contract, and it ships CombinatorialPurgedCV and WalkForward as first-class cross-validators in skfolio.model_selection. That collapses two requirements - portfolio optimization and finance-appropriate validation - into one dependency and eliminates the need for mlfinpy entirely. Import as: from skfolio.model_selection import CombinatorialPurgedCV, WalkForward"
},
{
"package": "vollib",
"version": "1.0.11",
"latest_release_date": "2026-06-01",
"license_spdx": "MIT",
"repo_url": "https://github.com/vollib/py_vollib",
"stars": 420,
"open_issues": 1,
"maintenance_status": "active",
"architecture_layer": "Options - BSM price / Greeks / IV",
"verified_via": "unverified - neither surviving source captured a single PyPI or GitHub API response (no JSON excerpts, no HTTP status codes, no per-record retrieval timestamps, no ETags), so every version string, release date, star count and open-issue count below is an unsourced assertion requiring independent re-verification via `pip index versions <name>` or https://pypi.org/pypi/<name>/json. The package SELECTION and architecture-layer assignments survive the audit; the metadata precision does not. Claims of `verified_via: PyPI API` in a source appendix are not reproduced here. | ROW NOTE: USE vollib, NOT py_vollib. Black, Black-Scholes and Black-Scholes-Merton analytic prices; the full standard Greek set; and implied volatility via Peter Jaeckel's 'Let's Be Rational' algorithm, which is essentially machine-precision and non-iterative - it matters when inverting thousands of quotes to build a surface. At USD 100 scale vollib alone is sufficient. Single-maintainer."
},
{
"package": "pyfeng",
"version": "0.5.0",
"latest_release_date": "2026-05-26",
"license_spdx": "GPL-2.0",
"repo_url": "https://github.com/PyFE/PyFENG",
"stars": 184,
"open_issues": 2,
"maintenance_status": "active",
"architecture_layer": "Options - SABR / Heston / rough vol",
"verified_via": "unverified - neither surviving source captured a single PyPI or GitHub API response (no JSON excerpts, no HTTP status codes, no per-record retrieval timestamps, no ETags), so every version string, release date, star count and open-issue count below is an unsourced assertion requiring independent re-verification via `pip index versions <name>` or https://pypi.org/pypi/<name>/json. The package SELECTION and architecture-layer assignments survive the audit; the metadata precision does not. Claims of `verified_via: PyPI API` in a source appendix are not reproduced here. | ROW NOTE: THE STRICTEST COPYLEFT IN THIS CLUSTER. Academic-grade and entirely appropriate for private research; do not distribute derived code without understanding the obligation. Add only when you need stochastic volatility."
},
{
"package": "QuantLib",
"version": "1.43",
"latest_release_date": "2026-07-14",
"license_spdx": "BSD-3-Clause",
"repo_url": "https://github.com/lballabio/QuantLib-SWIG",
"stars": 1900,
"open_issues": 35,
"maintenance_status": "active",
"architecture_layer": "Options - term structure / exotics",
"verified_via": "unverified - neither surviving source captured a single PyPI or GitHub API response (no JSON excerpts, no HTTP status codes, no per-record retrieval timestamps, no ETags), so every version string, release date, star count and open-issue count below is an unsourced assertion requiring independent re-verification via `pip index versions <name>` or https://pypi.org/pypi/<name>/json. The package SELECTION and architecture-layer assignments survive the audit; the metadata precision does not. Claims of `verified_via: PyPI API` in a source appendix are not reproduced here. | ROW NOTE: Version resolved to the higher (1.43 vs 1.35); repository resolved AGAINST the higher-version source, which presented lballabio as a MIRROR of a canonical quantlib/QuantLib repo, inverting the actual relationship. lballabio is upstream and QuantLib-SWIG is correct for the Python bindings. This is the one repository-attribution conflict the other source wins outright. Add only when you need a real term structure."
},
{
"package": "FinancePy",
"version": "1.0.1",
"latest_release_date": "2025-08-31",
"license_spdx": "GPL-3.0-or-later",
"repo_url": "https://github.com/domokane/FinancePy",
"stars": 3080,
"open_issues": 55,
"maintenance_status": "slowing",
"architecture_layer": "Options / rates / credit (optional)",
"verified_via": "unverified - neither surviving source captured a single PyPI or GitHub API response (no JSON excerpts, no HTTP status codes, no per-record retrieval timestamps, no ETags), so every version string, release date, star count and open-issue count below is an unsourced assertion requiring independent re-verification via `pip index versions <name>` or https://pypi.org/pypi/<name>/json. The package SELECTION and architecture-layer assignments survive the audit; the metadata precision does not. Claims of `verified_via: PyPI API` in a source appendix are not reproduced here. | ROW NOTE: Overkill for a pure-options track."
},
{
"package": "scores",
"version": "2.6.0",
"latest_release_date": "2026-07-17",
"license_spdx": "Apache-2.0",
"repo_url": "https://github.com/nci/scores",
"stars": 228,
"open_issues": 104,
"maintenance_status": "active",
"architecture_layer": "Evaluation - Brier / CRPS / PIT / Diebold-Mariano",
"verified_via": "unverified - neither surviving source captured a single PyPI or GitHub API response (no JSON excerpts, no HTTP status codes, no per-record retrieval timestamps, no ETags), so every version string, release date, star count and open-issue count below is an unsourced assertion requiring independent re-verification via `pip index versions <name>` or https://pypi.org/pypi/<name>/json. The package SELECTION and architecture-layer assignments survive the audit; the metadata precision does not. Claims of `verified_via: PyPI API` in a source appendix are not reproduced here. | ROW NOTE: THE LAYER THAT CONVERTS A TRADING EXPERIMENT INTO A MEASURABLE ONE, and the most comprehensive coverage of the meteorological metric set available in Python: Brier and threshold-Brier, CRPS, FIRM, SEEPS, MAE/MSE/RMSE, Kling-Gupta Efficiency, NSE, Flip-Flop Index, the Diebold-Mariano test, Fractions Skill Score, and isotonic regression for reliability diagrams. Both sources recommend it and both assign it to the evaluation layer - the cleanest cross-source agreement in the cluster on ROLE. Repository resolved to nci/scores (Australia's NCI); the competing nswbusiness/scores owner is flagged fabricated-looking by its own digest. CAUTION: the package organizes metrics under submodules, so read the module layout before writing top-level imports."
},
{
"package": "scoringrules",
"version": "0.11.0",
"latest_release_date": "2026-06-06",
"license_spdx": "Apache-2.0",
"repo_url": "https://github.com/frazane/scoringrules",
"stars": 97,
"open_issues": 17,
"maintenance_status": "active",
"architecture_layer": "Evaluation - fast multi-backend CRPS",
"verified_via": "unverified - neither surviving source captured a single PyPI or GitHub API response (no JSON excerpts, no HTTP status codes, no per-record retrieval timestamps, no ETags), so every version string, release date, star count and open-issue count below is an unsourced assertion requiring independent re-verification via `pip index versions <name>` or https://pypi.org/pypi/<name>/json. The package SELECTION and architecture-layer assignments survive the audit; the metadata precision does not. Claims of `verified_via: PyPI API` in a source appendix are not reproduced here. | ROW NOTE: CRPS, energy, variogram, interval and quantile scores across NumPy, JAX, PyTorch and TensorFlow backends. Take it when CRPS evaluation is in an inner loop and speed matters."
},
{
"package": "xskillscore",
"version": "0.0.29",
"latest_release_date": "2026-02-18",
"license_spdx": "Apache-2.0",
"repo_url": "https://github.com/xarray-contrib/xskillscore",
"stars": 242,
"open_issues": 52,
"maintenance_status": "slowing",
"architecture_layer": "Evaluation - xarray skill scores",
"verified_via": "unverified - neither surviving source captured a single PyPI or GitHub API response (no JSON excerpts, no HTTP status codes, no per-record retrieval timestamps, no ETags), so every version string, release date, star count and open-issue count below is an unsourced assertion requiring independent re-verification via `pip index versions <name>` or https://pypi.org/pypi/<name>/json. The package SELECTION and architecture-layer assignments survive the audit; the metadata precision does not. Claims of `verified_via: PyPI API` in a source appendix are not reproduced here. | ROW NOTE: The right tool only if forecasts live in xarray."
},
{
"package": "scikit-learn",
"version": "1.9.0",
"latest_release_date": "2026-06-02",
"license_spdx": "BSD-3-Clause",
"repo_url": "https://github.com/scikit-learn/scikit-learn",
"stars": 66849,
"open_issues": 2109,
"maintenance_status": "active",
"architecture_layer": "ML estimators + calibration + TimeSeriesSplit",
"verified_via": "unverified - neither surviving source captured a single PyPI or GitHub API response (no JSON excerpts, no HTTP status codes, no per-record retrieval timestamps, no ETags), so every version string, release date, star count and open-issue count below is an unsourced assertion requiring independent re-verification via `pip index versions <name>` or https://pypi.org/pypi/<name>/json. The package SELECTION and architecture-layer assignments survive the audit; the metadata precision does not. Claims of `verified_via: PyPI API` in a source appendix are not reproduced here. | ROW NOTE: Both sources agree on the version to within one day of release date - the strongest metadata corroboration anywhere in this cluster. IT DOES NOT IMPLEMENT PURGED OR EMBARGOED CROSS-VALIDATION, and this is the gap that destroys most retail ML backtests: overlapping labels leak across naive k-fold boundaries and produce out-of-sample collapse. Use skfolio's splitters, never sklearn.model_selection.KFold."
},
{
"package": "lightgbm",
"version": "4.7.0",
"latest_release_date": "2026-05-04",
"license_spdx": "MIT",
"repo_url": "https://github.com/microsoft/LightGBM",
"stars": 16500,
"open_issues": 340,
"maintenance_status": "active",
"architecture_layer": "ML - gradient boosting",
"verified_via": "unverified - neither surviving source captured a single PyPI or GitHub API response (no JSON excerpts, no HTTP status codes, no per-record retrieval timestamps, no ETags), so every version string, release date, star count and open-issue count below is an unsourced assertion requiring independent re-verification via `pip index versions <name>` or https://pypi.org/pypi/<name>/json. The package SELECTION and architecture-layer assignments survive the audit; the metadata precision does not. Claims of `verified_via: PyPI API` in a source appendix are not reproduced here. | ROW NOTE: SINGLE-SOURCE with unverifiable metadata; the other source omitted the entire gradient-boosting tier despite nominally covering it. Existence and maintenance status are not seriously in doubt, and a finance ML layer without a boosted-tree implementation is incomplete. Wrap in skfolio's CV splitters."
},
{
"package": "xgboost",
"version": "3.3.0",
"latest_release_date": "2026-06-20",
"license_spdx": "Apache-2.0",
"repo_url": "https://github.com/dmlc/xgboost",
"stars": 26100,
"open_issues": 480,
"maintenance_status": "active",
"architecture_layer": "ML - gradient boosting",
"verified_via": "unverified - neither surviving source captured a single PyPI or GitHub API response (no JSON excerpts, no HTTP status codes, no per-record retrieval timestamps, no ETags), so every version string, release date, star count and open-issue count below is an unsourced assertion requiring independent re-verification via `pip index versions <name>` or https://pypi.org/pypi/<name>/json. The package SELECTION and architecture-layer assignments survive the audit; the metadata precision does not. Claims of `verified_via: PyPI API` in a source appendix are not reproduced here. | ROW NOTE: Single-source; see lightgbm."
},
{
"package": "catboost",
"version": "1.2.8",
"latest_release_date": "2026-04-29",
"license_spdx": "Apache-2.0",
"repo_url": "https://github.com/catboost/catboost",
"stars": 8100,
"open_issues": 390,
"maintenance_status": "active",
"architecture_layer": "ML - categorical gradient boosting",
"verified_via": "unverified - neither surviving source captured a single PyPI or GitHub API response (no JSON excerpts, no HTTP status codes, no per-record retrieval timestamps, no ETags), so every version string, release date, star count and open-issue count below is an unsourced assertion requiring independent re-verification via `pip index versions <name>` or https://pypi.org/pypi/<name>/json. The package SELECTION and architecture-layer assignments survive the audit; the metadata precision does not. Claims of `verified_via: PyPI API` in a source appendix are not reproduced here. | ROW NOTE: Single-source; see lightgbm."
},
{
"package": "pydantic",
"version": "2.13.4",
"latest_release_date": "2026-05-06",
"license_spdx": "MIT",
"repo_url": "https://github.com/pydantic/pydantic",
"stars": 22400,
"open_issues": 210,
"maintenance_status": "active",
"architecture_layer": "Validation - records / config / contracts",
"verified_via": "unverified - neither surviving source captured a single PyPI or GitHub API response (no JSON excerpts, no HTTP status codes, no per-record retrieval timestamps, no ETags), so every version string, release date, star count and open-issue count below is an unsourced assertion requiring independent re-verification via `pip index versions <name>` or https://pypi.org/pypi/<name>/json. The package SELECTION and architecture-layer assignments survive the audit; the metadata precision does not. Claims of `verified_via: PyPI API` in a source appendix are not reproduced here. | ROW NOTE: Version conflict resolved to the higher; the competing claim again pairs a lower version with a later date. Star count merged in from the other source to fill an n/a cell. Validates records and contracts - API responses, configuration, trade records, model artifacts."
},
{
"package": "pandera",
"version": "0.32.1",
"latest_release_date": "2026-06-29",
"license_spdx": "MIT",
"repo_url": "https://github.com/unionai-oss/pandera",
"stars": 4413,
"open_issues": 448,
"maintenance_status": "active",
"architecture_layer": "Validation - DataFrame schemas",
"verified_via": "unverified - neither surviving source captured a single PyPI or GitHub API response (no JSON excerpts, no HTTP status codes, no per-record retrieval timestamps, no ETags), so every version string, release date, star count and open-issue count below is an unsourced assertion requiring independent re-verification via `pip index versions <name>` or https://pypi.org/pypi/<name>/json. The package SELECTION and architecture-layer assignments survive the audit; the metadata precision does not. Claims of `verified_via: PyPI API` in a source appendix are not reproduced here. | ROW NOTE: GATE EVERY INGESTED DATAFRAME BEHIND A pandera SCHEMA before it reaches the analytical pipeline. This is the operational form of the data-integrity requirements the backtesting analysis derives from the overfitting literature: a schema check that fires on a silently changed yfinance column layout is worth more than any amount of downstream defensive coding. Replaces great-expectations, whose open-source edition is now maintenance-only."
},
{
"package": "prefect",
"version": "3.8.1",
"latest_release_date": "2026-07-30",
"license_spdx": "Apache-2.0",
"repo_url": "https://github.com/PrefectHQ/prefect",
"stars": 23518,
"open_issues": 823,
"maintenance_status": "active",
"architecture_layer": "Orchestration",
"verified_via": "unverified - neither surviving source captured a single PyPI or GitHub API response (no JSON excerpts, no HTTP status codes, no per-record retrieval timestamps, no ETags), so every version string, release date, star count and open-issue count below is an unsourced assertion requiring independent re-verification via `pip index versions <name>` or https://pypi.org/pypi/<name>/json. The package SELECTION and architecture-layer assignments survive the audit; the metadata precision does not. Claims of `verified_via: PyPI API` in a source appendix are not reproduced here. | ROW NOTE: Version conflict resolved to the higher. Both sources independently select prefect for the reference architecture, but the deciding factor at this scale is which mental model fits - prefect wraps imperative Python functions, dagster models asset-centric DAGs. PICK ONE."
},
{
"package": "dagster",
"version": "1.13.16",
"latest_release_date": "2026-07-30",
"license_spdx": "Apache-2.0",
"repo_url": "https://github.com/dagster-io/dagster",
"stars": 14110,
"open_issues": 1801,
"maintenance_status": "active",
"architecture_layer": "Orchestration (alternative to Prefect)",
"verified_via": "unverified - neither surviving source captured a single PyPI or GitHub API response (no JSON excerpts, no HTTP status codes, no per-record retrieval timestamps, no ETags), so every version string, release date, star count and open-issue count below is an unsourced assertion requiring independent re-verification via `pip index versions <name>` or https://pypi.org/pypi/<name>/json. The package SELECTION and architecture-layer assignments survive the audit; the metadata precision does not. Claims of `verified_via: PyPI API` in a source appendix are not reproduced here. | ROW NOTE: The credible alternative. apache-airflow is excluded for weight, not health - disproportionate for a single-user stack."
},
{
"package": "mlflow",
"version": "3.15.0",
"latest_release_date": "2026-07-31",
"license_spdx": "Apache-2.0",
"repo_url": "https://github.com/mlflow/mlflow",
"stars": 27318,
"open_issues": 2091,
"maintenance_status": "active",
"architecture_layer": "Logging - experiment tracking",
"verified_via": "unverified - neither surviving source captured a single PyPI or GitHub API response (no JSON excerpts, no HTTP status codes, no per-record retrieval timestamps, no ETags), so every version string, release date, star count and open-issue count below is an unsourced assertion requiring independent re-verification via `pip index versions <name>` or https://pypi.org/pypi/<name>/json. The package SELECTION and architecture-layer assignments survive the audit; the metadata precision does not. Claims of `verified_via: PyPI API` in a source appendix are not reproduced here. | ROW NOTE: Two-source version agreement. The local filesystem backend is entirely sufficient here. LOG EVERY BACKTEST RUN - parameters, metrics, artifacts - because the only defensible output of a 90-day USD 100 experiment is a complete, honest record of what was tried. Log the losers as carefully as the winners; the epistemic salvage plan depends entirely on this layer being honest."
},
{
"package": "optuna",
"version": "4.9.0",
"latest_release_date": "2026-06-01",
"license_spdx": "MIT",
"repo_url": "https://github.com/optuna/optuna",
"stars": 14591,
"open_issues": 18,
"maintenance_status": "active",
"architecture_layer": "Logging - hyperparameter search",
"verified_via": "unverified - neither surviving source captured a single PyPI or GitHub API response (no JSON excerpts, no HTTP status codes, no per-record retrieval timestamps, no ETags), so every version string, release date, star count and open-issue count below is an unsourced assertion requiring independent re-verification via `pip index versions <name>` or https://pypi.org/pypi/<name>/json. The package SELECTION and architecture-layer assignments survive the audit; the metadata precision does not. Claims of `verified_via: PyPI API` in a source appendix are not reproduced here. | ROW NOTE: USE SPARINGLY AND INSIDE PURGED CV. An HPO loop over a short financial time series is a machine for manufacturing overfit Sharpe ratios, and every multiple-testing correction applies to every trial it runs."
}
],
"unmaintained_libraries": [
{
"package": "backtrader 1.9.78.123",
"last_release_date": "2023-04-19",
"reason_excluded": "~23.5 months dormant at the research date; repository untouched since 2024-08. THE REPUTATIONAL TRAP THIS SECTION EXISTS TO PREVENT: roughly 22,660 stars make it the most-recommended dead backtester in Python, ubiquitous in tutorials and video courses. No native asyncio, no modern broker WebSocket drivers. Treat it as a frozen codebase, not a maintained library. Sources disagree on the exact date (2023-04-19 vs 2023-04-08); the later is carried. Replace with backtesting.py, vectorbt or nautilus-trader."
},
{
"package": "zipline 1.4.1",
"last_release_date": "2020-10-05",
"reason_excluded": "Quantopian shut down in late 2020 and upstream has had no commits since. Replace with nautilus-trader."
},
{
"package": "zipline-reloaded 3.1.1",
"last_release_date": "2025-07-19",
"reason_excluded": "THE SINGLE MOST CONSEQUENTIAL CORRECTION IN THE PYTHON SECTION. One source RECOMMENDS it as 'the live replacement' for zipline and marks it active-slowing - while its own reported release date falls 378 days before its own 2026-08-01 research date, failing its own mechanically-stated 365-day recency rule. Its arithmetic contradicts its prose and it did not notice. The other source independently places it on its avoid list citing Cython compilation failures and a hard pandas<2.0 pin, which is independently disqualifying against a pandas 3.x substrate. Replace with nautilus-trader or backtesting.py."
},
{
"package": "pyalgotrade 0.20",
"last_release_date": "2018-08-21",
"reason_excluded": "~8 years dormant; non-functional on Python 3.10+; no Arrow/Polars integration. Persists in legacy tutorials. Replace with vectorbt or backtesting.py."
},
{
"package": "pybacktest 1.1.8",
"last_release_date": "2025-03-27",
"reason_excluded": "Technically inside the recency window, but the single 2025 release was the first since 2015, against 4 stars, 5 issues, no community and no documentation. RECENCY IS NECESSARY, NOT SUFFICIENT. Replace with backtesting.py."
},
{
"package": "pyfolio 0.9.2",
"last_release_date": "2019-04-15",
"reason_excluded": "Quantopian lineage, abandoned ~7.3 years. Deprecated pandas/empyrical calls now raise at runtime. Sources disagree on the date (2019-04-15 vs 2019-06-21). Replace with skfolio's Portfolio.summary()."
},
{
"package": "pyfolio-reloaded",
"last_release_date": null,
"reason_excluded": "The commonly-recommended community fork WAS NOT FOUND ON PyPI as of 2026-08-01 - git repository only. Do NOT treat it as a drop-in successor. No release date exists to record. Replace with skfolio."
},
{
"package": "empyrical 0.5.5",
"last_release_date": "2020-10-13",
"reason_excluded": "Same Quantopian lineage; ~5.8 years since release. Replace with skfolio or Riskfolio-Lib."
},
{
"package": "mlfinlab (Hudson & Thames)",
"last_release_date": null,
"reason_excluded": "NO PyPI RELEASE AT ALL, so no release date exists. The 4.9k-star public GitHub repository contains 11 commits, no tagged releases, and a README declaring 'all rights reserved'; the real code ships under a paid commercial license. It is the most-cited Lopez de Prado implementation and IT IS NOT OPEN-SOURCE SOFTWARE. Replace with skfolio.model_selection.CombinatorialPurgedCV."
},
{
"package": "mlfinpy 0.1.2",
"last_release_date": "2024-10-09",
"reason_excluded": "661 days before the research date - the only release since project inception; repository dormant since 2025-01-23; 1 open issue against 79 stars. Fails recency decisively. Replace with skfolio.model_selection.CombinatorialPurgedCV."
},
{
"package": "timeseriescv 0.2",
"last_release_date": "2018-09-07",
"reason_excluded": "~7.9 years dormant. Purged walk-forward CV now lives in skfolio.model_selection.WalkForward."
},
{
"package": "filterpy 1.4.5",
"last_release_date": "2018-10-10",
"reason_excluded": "~7.8 years dormant. The foundational Kalman/EKF reference, no longer maintained. Replace with statsmodels.tsa.statespace or pykalman."
},
{
"package": "properscoring 0.1",
"last_release_date": "2015-11-12",
"reason_excluded": "Over a decade dormant, and STILL THE TOP SEARCH RESULT for Python proper scoring rules - precisely the failure mode this list exists to prevent. Sources disagree on the date (2015-11-12 vs 2015-05-20). Note the internal contradiction in one source: properscoring appears on its own AVOID list while being simultaneously recommended in its imported-techniques table and its machine-readable appendix. Replace with scores, scoringrules or xskillscore - or implement Brier and CRPS directly, which is under 30 lines of NumPy each."
},
{
"package": "uncertainty-toolbox",
"last_release_date": null,
"reason_excluded": "NOT LOCATED ON PyPI, so no release date exists. A Google research project, unmaintained since 2021. Replace with the scores reliability-diagram functions."
},
{
"package": "pyro-ppl 1.9.1",
"last_release_date": "2024-06-02",
"reason_excluded": "~26 months at the research date; fails recency. Replace with numpyro, which shares its modeling idioms."
},
{
"package": "py_vollib 1.0.12",
"last_release_date": "2026-06-01",
"reason_excluded": "DEPRECATED BY ITS OWN MAINTAINER - a transitional alias depending on vollib for the implementation following a 2026-06-01 rebrand. Passes recency but is explicitly end-of-life. Note the anomaly: version 1.0.12 EXCEEDS the canonical vollib 1.0.11 it wraps, and both carry the identical release date, so at least one of the two version numbers is wrong. The direction of the rename is consistent across every mention and is what matters operationally. Replace with vollib."
},
{
"package": "py-vollib-vectorized 0.1.1",
"last_release_date": "2021-02-28",
"reason_excluded": "~5.4 years dormant; vollib now vectorizes natively."
},
{
"package": "opstrat",
"last_release_date": "2021-07-14",
"reason_excluded": "Unmaintained options-plotting script. Replace with QuantLib, scipy or vollib."
},
{
"package": "ffn",
"last_release_date": "2022-08-15",
"reason_excluded": "~4 years dormant. Replace with Riskfolio-Lib or skfolio."
},
{
"package": "pyflux 0.9.1",
"last_release_date": null,
"reason_excluded": "Last released in 2017 - ~9 years dormant. Only the year is stated by the source, so no full date is recorded rather than guessing a month and day. Replace with statsmodels or pymc."
},
{
"package": "ib_insync",
"last_release_date": null,
"reason_excluded": "Abandoned in 2024 after the original maintainer's death; the project was renamed and transferred. Only the year is stated by the source, so no full date is recorded. Replace with ib-async at ib-api-reloaded/ib_async."
},
{
"package": "great-expectations (open-source edition)",
"last_release_date": null,
"reason_excluded": "The project moved to a commercial 'GX Core' and the open-source edition is maintenance-only. No release date is stated by the source. Replace with pandera."
}
],
"data_sources": [
{
"name": "SEC EDGAR",
"base_url": "https://data.sec.gov",
"asset_classes": [
"US filings",
"XBRL company facts",
"submissions",
"full-text search"
],
"history_depth": "Filings ~1990s onward; XBRL coverage begins later and varies by filer",
"update_latency": "Real time on filing acceptance",
"rate_limit": "10 req/sec documented fair-access ceiling",
"auth_required": false,
"point_in_time": false,
"survivorship_bias_free": false,
"tos_restricts_automation": false,
"genuinely_free": true,
"known_defects": "PIT = PARTIAL, not yes: reconstructable if indexed by filing acceptance timestamp, but the raw facts API is not itself a PIT database - amendments and taxonomy changes require event-time filtering, and XBRL tagging is inconsistent across filers. SBF = PARTIAL: filings of delisted issuers are retained, but EDGAR supplies NO PRICED SECURITY MASTER, so a survivorship-bias-free equity universe cannot be constructed from it. One source marked both YES; the finer-grained reading is adopted. A descriptive User-Agent header is mandatory and non-compliant clients are throttled or blocked - one source received a 403 during its own research pass. No redistribution restriction on the data itself."
},
{
"name": "FRED",
"base_url": "https://api.stlouisfed.org",
"asset_classes": [
"US and international macro",
"rates",
"labor",
"prices"
],
"history_depth": "Series-specific; longest series extend into the early 20th century",
"update_latency": "On release; revisions arrive after release",
"rate_limit": "120 req/min",
"auth_required": true,
"point_in_time": false,
"survivorship_bias_free": false,
"tos_restricts_automation": false,
"genuinely_free": true,
"known_defects": "PIT = NO on default endpoints and YES ONLY when realtime_start / realtime_end / vintage_dates are used. The qualifier is load-bearing: any pipeline that reads FRED's default endpoints and stores one value per series has already destroyed its own point-in-time property. Store release_timestamp, observation_date, vintage_date and source_series_id. Terms permit broad public use; third-party series retain source restrictions; use the official API rather than scraping."
},
{
"name": "ALFRED",
"base_url": "https://alfred.stlouisfed.org",
"asset_classes": [
"Archived vintages of FRED macro and rates series"
],
"history_depth": "Series-dependent; often decades of vintages",
"update_latency": "Vintage snapshots at release",
"rate_limit": "Same as FRED (120 req/min)",
"auth_required": true,
"point_in_time": true,
"survivorship_bias_free": false,
"tos_restricts_automation": false,
"genuinely_free": true,
"known_defects": "THE ONLY FREE ROW IN THE ENTIRE 33-SOURCE TABLE WITH AN UNQUALIFIED point_in_time = YES. Applies only to series with archived vintages. Same FRED terms and source-series restrictions."
},
{
"name": "US Treasury",
"base_url": "https://fiscaldata.treasury.gov",
"asset_classes": [
"Par yield curves",
"bill rates",
"auctions",
"debt",
"receipts and outlays"
],
"history_depth": "Dataset-specific, often decades",
"update_latency": "Daily, monthly or event-driven",
"rate_limit": "No universal documented quota; paginate and cache",
"auth_required": false,
"point_in_time": false,
"survivorship_bias_free": false,
"tos_restricts_automation": false,
"genuinely_free": true,
"known_defects": "PIT = NO for revised series; record timestamps vary by dataset. US government data broadly reusable; third-party marks and dataset notices apply."
},
{
"name": "Bureau of Labor Statistics",
"base_url": "https://www.bls.gov/developers",
"asset_classes": [
"CPI",
"employment",
"wages",
"productivity",
"release calendar"
],
"history_depth": "Series-specific, usually decades",
"update_latency": "Scheduled release; revision policy varies",
"rate_limit": "Documented daily request and row limits - batch and cache",
"auth_required": false,
"point_in_time": false,
"survivorship_bias_free": false,
"tos_restricts_automation": false,
"genuinely_free": true,
"known_defects": "PIT = NO unless release vintages are stored. Registration raises limits but is not required for v1. Reusable with attribution; preserve release metadata."
},
{
"name": "Bureau of Economic Analysis",
"base_url": "https://apps.bea.gov/API",
"asset_classes": [
"GDP",
"personal income",
"trade",
"industry accounts"
],
"history_depth": "National-accounts series, often decades",
"update_latency": "Scheduled releases with revisions",
"rate_limit": "Quota tied to the API account",
"auth_required": true,
"point_in_time": false,
"survivorship_bias_free": false,
"tos_restricts_automation": false,
"genuinely_free": true,
"known_defects": "PIT = NO without vintage capture. Reusable; third-party inputs and trademarks may differ."
},
{
"name": "ECB Data Portal",
"base_url": "https://data.ecb.europa.eu",
"asset_classes": [
"Euro-area rates",
"FX",
"macro",
"banking and financial statistics"
],
"history_depth": "Dataset-specific, frequently decades",
"update_latency": "Release and event dependent; revisions occur",
"rate_limit": "No single documented universal quota",
"auth_required": false,
"point_in_time": false,
"survivorship_bias_free": false,
"tos_restricts_automation": false,
"genuinely_free": true,
"known_defects": "PIT = NO unless vintage/release metadata is retained. ECB legal notices and dataset-specific reuse terms apply."
},
{
"name": "Bank of England Interactive Database",
"base_url": "https://www.bankofengland.co.uk/boeapps/database",
"asset_classes": [
"UK rates",
"yield curves",
"macro and financial series"
],
"history_depth": "Dataset-specific, often decades",
"update_latency": "Daily or release-based",
"rate_limit": "None verified",
"auth_required": false,
"point_in_time": false,
"survivorship_bias_free": false,
"tos_restricts_automation": false,
"genuinely_free": true,
"known_defects": "PIT = NO without vintage storage. Bank terms and copyright notices apply."
},
{
"name": "Yahoo Finance / yfinance",
"base_url": "https://query2.finance.yahoo.com",
"asset_classes": [
"US and global equities",
"ETFs",
"FX",
"crypto",
"options chains",
"fundamentals",
"news"
],
"history_depth": "~30y daily, ~60d intraday; the provider guarantees no retention",
"update_latency": "Quotes near-real-time to 15-min delayed; historical-endpoint latency undocumented",
"rate_limit": "No published limit; ~2,000 req/hr observed unofficially (both sources agree no official limit exists)",
"auth_required": false,
"point_in_time": false,
"survivorship_bias_free": false,
"tos_restricts_automation": true,
"genuinely_free": false,
"known_defects": "THE SINGLE MOST-USED INGESTION SOURCE IN RETAIL QUANT AND THE ONE OPERATING FURTHEST OUTSIDE ITS PROVIDER'S TERMS. Terms contemplate personal/non-commercial use; automated extraction and redistribution are restricted; the interface is unofficial and IP bans are reported. Freemium, not genuinely free. Unannounced schema breaks, adjusted-price gaps, no delisted stocks, and a cookie/crumb handshake that varies and breaks. Corroborated across both surviving sources as the most common cause of survivorship bias in retail backtests. Neither report quotes the governing clause, so this is a legal-review item, not a settled fact."
},
{
"name": "Alpha Vantage",
"base_url": null,
"asset_classes": [
"Global equities/ETFs",
"splits and dividends",
"fundamentals",
"earnings calendar and estimates",
"options",
"news/sentiment",
"FX",
"crypto",
"commodities",
"macro"
],
"history_depth": "20+ years daily - BUT full daily history is a premium entitlement and the free tier is capped",
"update_latency": "Historical; delayed and real-time reserved to premium",
"rate_limit": "Free quota is plan-dependent and revised without notice",
"auth_required": true,
"point_in_time": false,
"survivorship_bias_free": false,
"tos_restricts_automation": true,
"genuinely_free": false,
"known_defects": "The '20+ years daily' figure must never be lifted without the premium qualifier - full daily history and most intraday access are premium, so the free key supports exploratory pulls, not broad universe ingestion. Its documented real-time and historical US options, put-call ratios and volume/open-interest ratios are ALL marked premium. Exchange-data policy and commercial entitlement restrictions on redistribution and derived use."
},
{
"name": "Polygon.io",
"base_url": null,
"asset_classes": [
"US equities/ETFs",
"options",
"futures",
"FX",
"crypto",
"corporate actions"
],
"history_depth": "Plan- and asset-dependent; the free tier is not a complete archive",
"update_latency": "Free tier delayed; latency is plan-dependent",
"rate_limit": "Plan-specific",
"auth_required": true,
"point_in_time": false,
"survivorship_bias_free": false,
"tos_restricts_automation": true,
"genuinely_free": false,
"known_defects": "Market-data licensing and plan terms restrict redistribution and commercial derived products. One source asserts a Polygon-to-'Massive' corporate rebrand on 2025-10-30 with a new canonical package; the asserted rebrand date is identical to the asserted release date of polygon-api-client 1.16.3 and the claim could not be corroborated."
},
{
"name": "Nasdaq Data Link (ex-Quandl)",
"base_url": null,
"asset_classes": [
"Macro",
"fundamentals",
"equities",
"futures",
"rates - dataset-specific"
],
"history_depth": "Dataset-specific; many free datasets have finite history",
"update_latency": "End-of-day or periodic",
"rate_limit": "Dataset-specific",
"auth_required": true,
"point_in_time": false,
"survivorship_bias_free": false,
"tos_restricts_automation": true,
"genuinely_free": false,
"known_defects": "Per-dataset license plus platform terms govern automated and derived commercial use."
},
{
"name": "Tiingo",
"base_url": null,
"asset_classes": [
"Equities/ETFs",
"fundamentals",
"news",
"crypto"
],
"history_depth": "Plan-dependent; free account limited",
"update_latency": "Delayed or end-of-day by feed",
"rate_limit": "Plan-dependent; no universal free limit",
"auth_required": true,
"point_in_time": false,
"survivorship_bias_free": false,
"tos_restricts_automation": true,
"genuinely_free": false,
"known_defects": "Terms plus exchange redistribution restrictions. A free token does not imply commercial rights."
},
{
"name": "IEX Cloud",
"base_url": null,
"asset_classes": [
"US equities",
"fundamentals",
"corporate actions"
],
"history_depth": "Unresolved",
"update_latency": "Unresolved",
"rate_limit": "Unresolved",
"auth_required": true,
"point_in_time": false,
"survivorship_bias_free": false,
"tos_restricts_automation": true,
"genuinely_free": false,
"known_defects": "DO NOT ADOPT WITHOUT VERIFYING THE PRODUCT STILL EXISTS. One source hedges service status throughout; the other does not cover it. The suggestion that IEX Cloud retired its data products comes from a digest author's own annotation rather than from any source report and is therefore not carried as a documented fact - the row is downgraded rather than deleted so the hedge is not silently propagated as a live recommendation."
},
{
"name": "Marketstack",
"base_url": null,
"asset_classes": [
"Global equities",
"EOD and intraday",
"corporate actions by plan"
],
"history_depth": "Free plan limited to recent history and request volume",
"update_latency": "Free tier delayed; plan-specific",
"rate_limit": "Free-plan request quota is pricing-dependent",
"auth_required": true,
"point_in_time": false,
"survivorship_bias_free": false,
"tos_restricts_automation": true,
"genuinely_free": false,
"known_defects": "Commercial and redistribution rights are plan-dependent."
},
{
"name": "EOD Historical Data (EODHD)",
"base_url": null,
"asset_classes": [
"Global equities/ETFs",
"corporate actions",
"fundamentals",
"calendars",
"options by plan"
],
"history_depth": "Plan-dependent; free and demo access limited",
"update_latency": "EOD or delayed; plan-specific",
"rate_limit": "Plan-specific",
"auth_required": true,
"point_in_time": false,
"survivorship_bias_free": false,
"tos_restricts_automation": true,
"genuinely_free": false,
"known_defects": "Terms and dataset entitlements restrict automated redistribution and commercial derived use."
},
{
"name": "Stooq",
"base_url": null,
"asset_classes": [
"Equities",
"indices",
"FX",
"futures",
"ETFs - daily history"
],
"history_depth": "Broad daily archives, instrument-dependent; NO SECURITY MASTER",
"update_latency": "End-of-day / delayed",
"rate_limit": "No official API SLA or published limit located",
"auth_required": false,
"point_in_time": false,
"survivorship_bias_free": false,
"tos_restricts_automation": true,
"genuinely_free": false,
"known_defects": "The access pattern is a scrape. Automated-download permission must be established before polling or redistributing. Free for limited use, terms-sensitive. Classified as an interface likely to break; carry an identified replacement feed from day one."
},
{
"name": "Financial Modeling Prep (FMP)",
"base_url": null,
"asset_classes": [
"Fundamentals",
"analyst estimates",
"earnings calendars"
],
"history_depth": "Not established by any surveyed report",
"update_latency": "Not established",
"rate_limit": "Not established",
"auth_required": true,
"point_in_time": false,
"survivorship_bias_free": false,
"tos_restricts_automation": false,
"genuinely_free": false,
"known_defects": "NOT ESTABLISHED on history depth, latency, rate limit, TOS or free status by any surveyed report. tos_restricts_automation is recorded false because no source established that it restricts, not because any source established that it permits. Freemium status itself is not established."
},
{
"name": "CoinGecko",
"base_url": null,
"asset_classes": [
"Crypto prices",
"markets",
"exchanges",
"metadata"
],
"history_depth": "Endpoint- and plan-dependent; free history limited versus paid",
"update_latency": "Free public API delayed or rate-limited",
"rate_limit": "Plan-specific; do not assume a permanent free number",
"auth_required": true,
"point_in_time": false,
"survivorship_bias_free": false,
"tos_restricts_automation": true,
"genuinely_free": false,
"known_defects": "Terms distinguish personal/free from commercial use and redistribution. Aggregation defects: heterogeneous clocks, venue outages, symbol-mapping drift and wash-trading exposure."
},
{
"name": "CryptoCompare",
"base_url": null,
"asset_classes": [
"Crypto spot",
"OHLCV",
"trades",
"news",
"some derivatives"
],
"history_depth": "Asset-, exchange- and endpoint-dependent",
"update_latency": "Near-real-time or delayed by endpoint",
"rate_limit": "Account- and endpoint-dependent",
"auth_required": true,
"point_in_time": false,
"survivorship_bias_free": false,
"tos_restricts_automation": true,
"genuinely_free": false,
"known_defects": "Terms plus exchange-source rights restrict redistribution and commercial use. Same aggregation defects as CoinGecko."
},
{
"name": "Kaiko",
"base_url": null,
"asset_classes": [
"Institutional crypto spot",
"derivatives",
"order books"
],
"history_depth": "Not free for production; trial or demo may be offered",
"update_latency": "Tick and order-book latency depends on paid plan",
"rate_limit": "Contract-specific",
"auth_required": true,
"point_in_time": false,
"survivorship_bias_free": false,
"tos_restricts_automation": true,
"genuinely_free": false,
"known_defects": "PIT and SBF both UNKNOWN - not established either way by any source. Commercial license required; not free."
},
{
"name": "Deribit",
"base_url": "https://docs.deribit.com",
"asset_classes": [
"BTC/ETH options",
"futures",
"order books",
"trades",
"instrument metadata"
],
"history_depth": "Exchange- and endpoint-dependent",
"update_latency": "Real-time REST and WebSocket",
"rate_limit": "Exchange-specific published limits - check before polling",
"auth_required": false,
"point_in_time": false,
"survivorship_bias_free": false,
"tos_restricts_automation": true,
"genuinely_free": false,
"known_defects": "Free for public endpoints but terms-sensitive: exchange terms control automated access and redistribution, and public does not mean unrestricted commercial reuse. Exchange-specific crypto options only - NOT a US equity-options substitute."
},
{
"name": "Binance",
"base_url": "https://api.binance.com",
"asset_classes": [
"Crypto spot",
"crypto futures"
],
"history_depth": "2017-present",
"update_latency": "Real-time REST and WebSocket",
"rate_limit": "1,200 req/min",
"auth_required": false,
"point_in_time": false,
"survivorship_bias_free": false,
"tos_restricts_automation": false,
"genuinely_free": false,
"known_defects": "PIT = PARTIAL: printed trades are point-in-time by construction, but coverage is venue-scoped. SBF = NO: venue-scoped, and delisted pairs are not reliably retained - exchange feeds cover only their own venue and their own listed-instrument lifecycle, which is precisely why they cannot be survivorship-bias-free at the asset-universe level. One source marked both YES; downgraded by applying the other source's stated principle. GEOFENCED FOR US IPs - Binance.US required. Free for public endpoints (scope-qualified); automated access explicitly supported, redistribution governed by exchange terms."
},
{
"name": "Kraken",
"base_url": "https://api.kraken.com",
"asset_classes": [
"Crypto spot"
],
"history_depth": "Inception-present",
"update_latency": "Real-time REST and WebSocket",
"rate_limit": "1 req/sec public; REST OHLC returns max 720 bars per request (requires pagination)",
"auth_required": false,
"point_in_time": false,
"survivorship_bias_free": false,
"tos_restricts_automation": false,
"genuinely_free": false,
"known_defects": "PIT = PARTIAL, same venue-scoped caveat as Binance. SBF = NO, venue-scoped. Free for public endpoints (scope-qualified); automated access explicitly supported, redistribution governed by exchange terms."
},
{
"name": "Kalshi",
"base_url": "https://docs.kalshi.com",
"asset_classes": [
"CFTC-regulated event contracts - economic, political, climate, company",
"order books",
"trades",
"settlements"
],
"history_depth": "2021-present; market-history retention is endpoint-specific with NO BLANKET GUARANTEE",
"update_latency": "Real-time REST and WebSocket; settlement after official resolution",
"rate_limit": "Basic tier token bucket: 200 read + 100 write tokens/sec against a default request cost of 10 tokens, i.e. ~20 read and ~10 write req/sec. 429 responses OMIT Retry-After, so client backoff must be self-managed.",
"auth_required": true,
"point_in_time": false,
"survivorship_bias_free": false,
"tos_restricts_automation": false,
"genuinely_free": false,
"known_defects": "PIT = NO and SBF = NO: one source marked both YES but offered no evidence of a documented historical archive; the other cites the documentation and prescribes LOCAL ARCHIVING as the remedy. ARCHIVE EVERY MARKET, SERIES, CLOSE TIME, SETTLEMENT VALUE AND RULE TEXT LOCALLY AT CAPTURE TIME. tos_restricts_automation is false because automated ACCESS is supported through a documented API - but redistribution and derived commercial products ARE restricted by terms, and account eligibility applies. A request succeeding is not a license. Free apart from trading fees (scope-qualified). The rate-limit figure is the single load-bearing hard number most worth independent verification: the two sources plausibly read the same documentation page, so their agreement is corroboration rather than independent confirmation. No first-class Python client exists on PyPI - a thin httpx client must be hand-rolled."
},
{
"name": "ForecastEx (via Interactive Brokers)",
"base_url": "https://www.interactivebrokers.com",
"asset_classes": [
"CFTC-regulated event contracts"
],
"history_depth": "Inception-present; NO FREE HISTORICAL L2/L3 ORDER-BOOK DEPTH",
"update_latency": "Real-time via TWS/Gateway",
"rate_limit": "50 req/sec",
"auth_required": true,
"point_in_time": false,
"survivorship_bias_free": false,
"tos_restricts_automation": false,
"genuinely_free": false,
"known_defects": "PIT = NO and SBF = NO, resolved USING THE ASSERTING SOURCE'S OWN TEXT: its Table F marked both YES while the same report states elsewhere that no free historical L2/L3 order-book depth exists for this venue, and a venue whose depth history does not exist cannot support point-in-time order-book reconstruction. An internal contradiction within a single report is stronger evidence than cross-report disagreement. Requires a LOCALLY RUNNING IBKR TWS or Gateway process, which makes it an infrastructure dependency rather than a plain HTTP endpoint. Free with a funded IBKR account (scope-qualified); IBKR market-data terms apply."
},
{
"name": "Polymarket",
"base_url": null,
"asset_classes": [
"Prediction-market prices",
"trades",
"order books",
"settlements"
],
"history_depth": "Market-specific; no documented universal archive guarantee",
"update_latency": "Near-real-time; endpoint behavior changes without notice",
"rate_limit": "No stable published limit verified",
"auth_required": false,
"point_in_time": false,
"survivorship_bias_free": false,
"tos_restricts_automation": true,
"genuinely_free": false,
"known_defects": "Terms impose geographic, account, automated-use and IP restrictions; US participation requires separate legal verification. DO NOT INFER LEGALITY OR UNRESTRICTED AUTOMATED RIGHTS FROM THE EXISTENCE OF PUBLIC JSON. Undocumented endpoints and third-party clients are classified as likely to break. The three source reports are irreconcilable on US legal status and the dispute is routed to the regulatory analysis rather than resolved in a data table."
},
{
"name": "GDELT",
"base_url": "https://www.gdeltproject.org",
"asset_classes": [
"Global news events",
"tone/sentiment",
"entity extraction"
],
"history_depth": "Multi-decade event corpus",
"update_latency": "Near-real-time updates",
"rate_limit": "Not established",
"auth_required": false,
"point_in_time": false,
"survivorship_bias_free": false,
"tos_restricts_automation": false,
"genuinely_free": true,
"known_defects": "An EVENT-EXTRACTION CORPUS, NOT A LICENSED ARTICLE-TEXT FEED. Known defects requiring validation before any signal is derived: source duplication, language imbalance, timestamp ambiguity, entity-resolution errors. No source establishes complete historical news survivorship, stable article-text licensing, or bias-free sentiment labels - treat every sentiment field as a vendor-derived feature, never as ground truth."
},
{
"name": "Econoday / Trading Economics",
"base_url": null,
"asset_classes": [
"Economic-release calendars with consensus and actuals"
],
"history_depth": "Historical calendar depth is a paid, plan-gated feature",
"update_latency": "Event-time updates",
"rate_limit": "Plan-specific for API access; free web access is not an API license",
"auth_required": true,
"point_in_time": false,
"survivorship_bias_free": false,
"tos_restricts_automation": true,
"genuinely_free": false,
"known_defects": "A CALENDAR WITHOUT ARCHIVED 'WHAT WAS KNOWN WHEN' CONSENSUS CANNOT SUPPORT AN EVENT-SURPRISE BACKTEST. Knowing that CPI printed on a date tells you nothing about the surprise unless you also stored the forecast that existed before the print. Automated scraping and commercial reuse restricted by provider terms."
},
{
"name": "Investing.com",
"base_url": null,
"asset_classes": [
"Release calendars",
"prices",
"news"
],
"history_depth": "No authoritative archive guarantee",
"update_latency": "Web updates",
"rate_limit": "NO OFFICIAL PUBLIC API",
"auth_required": false,
"point_in_time": false,
"survivorship_bias_free": false,
"tos_restricts_automation": true,
"genuinely_free": false,
"known_defects": "Free web content, NOT a stable free API. Automated scraping and redistribution restricted; unofficial clients break. Calendar scraping is classified as an interface likely to break."
},
{
"name": "CRSP",
"base_url": null,
"asset_classes": [
"US equity prices",
"delisting returns",
"historical index constituents"
],
"history_depth": "Not stated by any surveyed report",
"update_latency": "n/a",
"rate_limit": "n/a",
"auth_required": true,
"point_in_time": true,
"survivorship_bias_free": true,
"tos_restricts_automation": true,
"genuinely_free": false,
"known_defects": "THE REFERENCE STANDARD for both point-in-time and survivorship-bias-free equity data, and one of only three rows in the table with an unqualified survivorship_bias_free = yes - all three of which are PAID. Institutional subscription; redistribution prohibited. Included in a free-data table on purpose: it is the benchmark the free universe is measured against, and its absence from the budget is a material limitation rather than a reason to substitute Yahoo data silently. Paid PIT equity data runs USD 100-500/month minimum - for a USD 100 experiment, the data required to make the backtest honest costs more than the capital at risk, every month."
},
{
"name": "Compustat / WRDS",
"base_url": null,
"asset_classes": [
"Fundamentals",
"point-in-time fundamentals products"
],
"history_depth": "Not stated by any surveyed report",
"update_latency": "n/a",
"rate_limit": "n/a",
"auth_required": true,
"point_in_time": true,
"survivorship_bias_free": true,
"tos_restricts_automation": true,
"genuinely_free": false,
"known_defects": "Reference standard for point-in-time fundamentals. Academic/institutional license. Same benchmark role as CRSP."
},
{
"name": "NYSE TAQ",
"base_url": null,
"asset_classes": [
"US trade and quote tick data"
],
"history_depth": "Not stated by any surveyed report",
"update_latency": "n/a",
"rate_limit": "n/a",
"auth_required": true,
"point_in_time": true,
"survivorship_bias_free": false,
"tos_restricts_automation": true,
"genuinely_free": false,
"known_defects": "Point-in-time by construction. survivorship_bias_free is recorded false because the source table marks it n/a rather than yes - the property is not applicable to a tick archive, and false here means 'not established as yes', not 'established as no'. Commercial license; subscription required."
}
],
"regulatory_requirements": [
{
"authority": "U.S. Securities and Exchange Commission",
"rule_citation": "Securities Exchange Act of 1934 sec. 15, 15 U.S.C. sec. 78o",
"applies_to": [
"broker-dealers"
],
"requirement": "Broker-dealers must register with the Commission.",
"consequence_at_100usd": "Falls entirely on the intermediary. A FINRA-member, SEC-registered broker-dealer is the necessary counterparty to any equity, ETF or listed-option position; the investor's own registration obligation is nil.",
"status": "settled",
"source_url": "https://www.law.cornell.edu/uscode/text/15/78o",
"retrieved": "2026-08-01"
},
{
"authority": "U.S. Securities and Exchange Commission",
"rule_citation": "Regulation Best Interest, 17 CFR sec. 240.15l-1",
"applies_to": [
"broker-dealers",
"retail equity and option accounts"
],
"requirement": "A broker-dealer must act in a retail customer's best interest when making a recommendation.",
"consequence_at_100usd": "Negligible. Reg BI attaches to RECOMMENDATIONS, and a purely self-directed account receives none. It does not bar a cash account and imposes no capital gate. Single-sourced.",
"status": "settled",
"source_url": "https://www.ecfr.gov/current/title-17/chapter-II/part-240/section-240.15l-1",
"retrieved": "2026-08-01"
},
{
"authority": "U.S. Securities and Exchange Commission",
"rule_citation": "SEC Rule 15c6-1, 17 CFR sec. 240.15c6-1 (amendment eff. 2024-05-28)",
"applies_to": [
"all cash accounts regardless of size",
"US securities settlement"
],
"requirement": "Standard U.S. securities settlement is T+1.",
"consequence_at_100usd": "Unsettled sale proceeds cannot fund the next purchase. Roughly ONE ROUND TRIP PER TWO BUSINESS DAYS on a single security, or ~30 round trips across a 90-day window under best-case timing; daily turnover capped at the stake. One source asserted T+2 three times and built its turnover analysis on it - stale by roughly two years, and the correction approximately doubles the achievable round-trip frequency.",
"status": "settled",
"source_url": "https://www.sec.gov/rules/final/34-96930.pdf",
"retrieved": "2026-08-01"
},
{
"authority": "FINRA",
"rule_citation": "FINRA Rule 2090 (Know Your Customer)",
"applies_to": [
"member firms",
"all account openings"
],
"requirement": "Member firms must use reasonable diligence to know the essential facts of every customer. Account opening requires SSN or ITIN, current address, employment, a financial profile and a risk-tolerance disclosure.",
"consequence_at_100usd": "Binds on the clock, not on the capital. Online onboarding runs 1-5 business days; manual review 1-4 weeks. The 90-day clock starts at tradeability, not at application.",
"status": "settled",
"source_url": "https://www.finra.org/rules-guidance/rulebooks/finra-rules/2090",
"retrieved": "2026-08-01"
},
{
"authority": "FinCEN (Bank Secrecy Act)",
"rule_citation": "Customer Identification Program - EXACT SECTION NUMBER UNRESOLVED. One source cites 17 CFR sec. 1010.230, which is wrong on its face (BSA CIP rules live in 31 CFR, administered by FinCEN, not in 17 CFR); the other cites 31 CFR sec. 1020.220, which is the BANKS subpart. The broker-dealer subpart is Part 1023 and no source cites it directly.",
"applies_to": [
"broker-dealers",
"all account openings"
],
"requirement": "A written Customer Identification Program with identity verification at account opening.",
"consequence_at_100usd": "The REQUIREMENT is settled and unanimous across sources; the section number that would let a reader look it up is not established by any of the three reports. Both sources invoke the USA PATRIOT Act with no section number; that reference was dropped. Initial ACH holds of 3-5 business days consume ~5.5% of the 90-day clock.",
"status": "unknown",
"source_url": "https://www.ecfr.gov/current/title-31/subtitle-B/chapter-X",
"retrieved": "2026-08-01"
},
{
"authority": "FINRA",
"rule_citation": "FINRA Rule 2111 (Suitability); FINRA Rule 3110 (Supervision)",
"applies_to": [
"member firms making recommendations"
],
"requirement": "A firm-level obligation to determine that a recommendation is suitable to the customer's investment profile.",
"consequence_at_100usd": "A purely self-directed, execution-only account receives no recommendation and therefore does not trigger the suitability obligation at all. The material exception: option and margin approvals inherently involve firm-level suitability review, because the firm must affirmatively approve the account for those privileges - self-direction does not bypass that gate. (A pincite to Rule 2111.05 offered by one source for the self-directed carve-out is uncorroborated and downgraded; the conclusion flows from the ABSENCE of a recommendation, not from an express exclusion.)",
"status": "settled",
"source_url": "https://www.finra.org/rules-guidance/rulebooks/finra-rules/2111",
"retrieved": "2026-08-01"
},
{
"authority": "FINRA / OCC / broker policy",
"rule_citation": "FINRA Rule 2360 (options account approval and firm diligence); Levels 1-4 are FIRM-SET INDUSTRY CONVENTION, not a FINRA-codified schedule",
"applies_to": [
"options accounts"
],
"requirement": "The firm must approve the account for options privileges. Level 1: covered calls, cash-secured puts. Level 2: buying calls and puts. Level 3: spreads, uncovered writing, married puts - most firms require a margin account. Level 4: uncovered/naked writing under strict Reg T or portfolio-margin requirements.",
"consequence_at_100usd": "LEVEL 2 IS THE PRACTICAL CEILING at USD 100 - long calls and puts only. The constraint is arithmetic rather than discretionary: spreads generally require a margin account and margin requires USD 2,000. Brokers additionally impose ~30 days of account seasoning for Level 3 and ~60 days for Level 4, consuming 33%-67% of a 90-day window before the strategy is even available. One source presented the ladder as codified at Rule 2360(b)(11)-(12); presenting it as a regulatory requirement would be wrong.",
"status": "contested",
"source_url": "https://www.finra.org/rules-guidance/rulebooks/finra-rules/2360",
"retrieved": "2026-08-01"
},
{
"authority": "FINRA / SEC",
"rule_citation": "FINRA Rule 4210(f)(8)(B) (Pattern Day Trader); NYSE-legacy analogue Rule 2520",
"applies_to": [
"margin accounts"
],
"requirement": "A pattern day trader - four or more day trades in five business days - must maintain USD 25,000 minimum equity.",
"consequence_at_100usd": "DOES NOT BIND. The PDT regime governs MARGIN accounts, and a USD 100 account must be a cash account, so it is outside the regime under either version of the rule. One source reports the rule rescinded effective 2026-06-04 and replaced by intraday margin standards (citing FINRA Notice 26-10, SEC approval 2026-04-14 at 91 FR 20731, broker phase-in through 2027-10-20), but that claim reaches the merge only through a tertiary chain with no primary notice retrieved, and the same source concedes some brokers may still enforce the USD 25,000 floor. Operationally moot either way. Separately, one source cited a nonexistent 'SEC Rule 2222' as the PDT authority and inverted the rule's logic; that citation was dropped.",
"status": "unsettled",
"source_url": "https://www.finra.org/rules-guidance/notices/26-10",
"retrieved": "2026-08-01"
},
{
"authority": "Federal Reserve Board / FINRA",
"rule_citation": "Regulation T, 12 CFR Part 220 (50% initial margin); FINRA Rule 4210(b)(4) (USD 2,000 minimum equity to open a margin account)",
"applies_to": [
"margin accounts"
],
"requirement": "Initial margin of 50% of the purchase price of marginable equity securities, and USD 2,000 minimum equity to open a margin account.",
"consequence_at_100usd": "THE HARDEST CONSTRAINT IN THE SECTION AND THE ONE POINT WHERE EVERY SOURCE AGREES WITHOUT QUALIFICATION. A USD 100 account is 5% of the way to the floor. Margin is unavailable, which forecloses spreads, uncovered writing and short selling, and forces a cash account.",
"status": "settled",
"source_url": "https://www.ecfr.gov/current/title-12/chapter-II/subchapter-A/part-220",
"retrieved": "2026-08-01"
},
{
"authority": "Federal Reserve Board (not FINRA)",
"rule_citation": "Regulation T, 12 CFR Part 220 - free-riding / good-faith violation; cited at part level because both sources' pinpoint attributions are wrong (one attributes it to FINRA Rule 2210, Communications with the Public; the other lists FINRA as the authority for a Federal Reserve Board regulation)",
"applies_to": [
"cash accounts"
],
"requirement": "In a cash account every transaction must be fully paid for. Selling a security before the purchase that acquired it has settled is a free-riding violation.",
"consequence_at_100usd": "A 90-day cash-up-front account restriction - which for a 90-day experiment means a single settlement mistake ends the window. TRIGGER COUNT UNRESOLVED: one violation (one source, one row) versus three good-faith violations in twelve months (the other source, and the first source's adjacent row). The three-in-twelve-months count is single-sourced to investor-education material rather than rule text. Directly caps maximum trade count.",
"status": "contested",
"source_url": "https://www.sec.gov/investor/alerts/cashaccounts.pdf",
"retrieved": "2026-08-01"
},
{
"authority": "Commodity Futures Trading Commission",
"rule_citation": "Commodity Exchange Act sec. 2(a)(1)(A), 7 U.S.C. sec. 2(a)(1)(A); swaps under CEA sec. 1a(47) and sec. 2(h)",
"applies_to": [
"futures contracts on commodities",
"event contracts",
"spot crypto derivatives"
],
"requirement": "The CFTC holds exclusive jurisdiction over futures contracts on commodities.",
"consequence_at_100usd": "Event contracts, futures, spot crypto and crypto derivatives on CFTC-registered venues are not directly SEC-jurisdictional; equity ETFs, listed options, security-based swaps and security futures are. One source labelled this Title 7 Commodity Exchange Act provision 'Securities Act sec. 2(a)(1)(A)'; the correct labelling is adopted.",
"status": "settled",
"source_url": "https://www.law.cornell.edu/uscode/text/7/2",
"retrieved": "2026-08-01"
},
{
"authority": "Commodity Futures Trading Commission",
"rule_citation": "CEA sec. 5c(c)(5)(C); 17 CFR sec. 40.11 (sec. 40.11(a)(1) carries the prohibition; review-and-approval mechanism runs 90 days)",
"applies_to": [
"designated contract markets",
"event contracts",
"prediction markets"
],
"requirement": "DCMs may not list event contracts involving terrorism, assassination, war, gaming or unlawful activity, subject to Commission review.",
"consequence_at_100usd": "This is one of the few citations in the entire regulatory analysis that two sources pin to the same authority with the same subsection, and it is accordingly the most reliable regulatory citation in the merged report. It governs whether the venue may offer the contract at all; it does not resolve the separate state-law question below.",
"status": "settled",
"source_url": "https://www.law.cornell.edu/cfr/text/17/40.11",
"retrieved": "2026-08-01"
},
{
"authority": "Commodity Futures Trading Commission",
"rule_citation": "Notice of proposed rulemaking at 91 FR 35806 (June 12, 2026), release 9249-26 - proposed 'Reg 40.11 Appendix F'",
"applies_to": [
"designated contract markets",
"event contracts"
],
"requirement": "A proposed framework for evaluating whether an event contract involves an enumerated activity or is contrary to the public interest.",
"consequence_at_100usd": "NOT ADOPTED - in public-comment phase at the research date, and a rule that is not adopted binds no one. Recorded because a reader planning a 90-day window should know the framework governing event-contract legality was actively in flux. Single-sourced; final text and adoption unknown.",
"status": "unsettled",
"source_url": "https://www.federalregister.gov/",
"retrieved": "2026-08-01"
},
{
"authority": "Commodity Futures Trading Commission",
"rule_citation": "CFTC release series 92xx-26 (including 9240-26 of May 29, 2026 approving KalshiEX BTCPERP, and Staff Letter 26-22 / release 9273-26 of July 24, 2026)",
"applies_to": [
"designated contract markets",
"KalshiEX BTCPERP"
],
"requirement": "Approval of KalshiEX's BTCPERP contract, classified by the Commission as a futures contract, and staff guidance advising DCMs to submit narrowly tailored rather than template self-certifications.",
"consequence_at_100usd": "THE ENTIRE 92xx-26 RELEASE-NUMBER SERIES IS UNVERIFIED against primary source, and the asserting source's own digest flags one release in the series as appearing in its bibliography while supporting nothing in its body - a hallucination signal. This matters because release 9240-26 anchors the Section 1256 tax conclusion for BTCPERP.",
"status": "unknown",
"source_url": "https://www.cftc.gov/PressRoom/PressReleases",
"retrieved": "2026-08-01"
},
{
"authority": "Commodity Futures Trading Commission",
"rule_citation": "CFTC designation as a contract market under CEA sec. 5 - KalshiEX LLC (Nov 2020); Polymarket via the QCEX acquisition with an Amended Order of Designation reported Nov 2025; ForecastEx LLC designation order NOT RETRIEVED by any source",
"applies_to": [
"KalshiEX LLC",
"Polymarket",
"ForecastEx LLC"
],
"requirement": "A venue offering event contracts must hold CFTC designation as a contract market.",
"consequence_at_100usd": "Kalshi's designation is the best-supported of the three but is sourced to Wikipedia and an unretrieved CFTC DCM list. Polymarket's designation carries no order number and no release number despite the same source citing six numbered releases elsewhere, so its federal authorization - which that source labels its most settled headline finding - rests on its weakest citation. FORECASTEX'S REGISTRATION STATUS IS UNVERIFIED, and a venue whose registration is unverified should not be carrying tax-treatment rows.",
"status": "unknown",
"source_url": "https://www.cftc.gov/IndustryOversight/TradingOrganizations/DCMs/index.htm",
"retrieved": "2026-08-01"
},
{
"authority": "Kalshi (DCM exchange rule)",
"rule_citation": "Kalshi fee schedule, July 7, 2026",
"applies_to": [
"every Kalshi trade"
],
"requirement": "Taker fee of ceil(0.07 x N x P x (1-P)) per side, rounded up to the cent.",
"consequence_at_100usd": "On a fixed stake this collapses to 7 x (1-P) dollars per side per USD 100, MONOTONICALLY DECREASING IN P: 7.0% of stake round trip at P = 0.50, 13.3% at P = 0.05, 1.4% at P = 0.90. Ceiling rounding alone costs 10% on a single USD 0.10 contract. The claim that the fee is 'U-shaped, so seek the tails' is refuted by the sources' own arithmetic once the denominator is the stake rather than the contract. Whether a settlement-side fee exists is established by no source, a ~2x uncertainty on every Kalshi figure. The competing 0.0175 coefficient offered by one source was self-tagged as derived from a help-center article rather than read from the fee schedule, and loses on both majority and quality; the correction scales every Kalshi friction figure by 4x.",
"status": "settled",
"source_url": "https://kalshi.com/docs/kalshi-fee-schedule.pdf",
"retrieved": "2026-08-01"
},
{
"authority": "Massachusetts Securities Division (Secretary of the Commonwealth)",
"rule_citation": "M.G.L. c. 110A (Massachusetts Uniform Securities Act)",
"applies_to": [
"securities offered or sold in Massachusetts",
"investment advisers"
],
"requirement": "The Massachusetts securities statute. Chapter 110A is adopted over one source's c. 110H (used ~14 times, alongside 'c. 110' and 'c. 110 sec. 410' in the same passage) because it is internally consistent and corroborated by a reported appellate decision under the same chapter; that source additionally supplied FEDERAL URLs as the source for its Massachusetts rows, an admission that no state authority was retrieved.",
"consequence_at_100usd": "No filing obligation for a principal trading their own USD 100. Section-level pincites within c. 110A from the source with the chapter defect are not propagated.",
"status": "settled",
"source_url": "https://malegislature.gov/Laws/GeneralLaws/PartI/TitleXV/Chapter110A",
"retrieved": null
},
{
"authority": "Massachusetts Securities Division (Secretary of the Commonwealth)",
"rule_citation": "950 CMR 12.207; Robinhood Financial LLC v. Secretary of the Commonwealth, 492 Mass. 696 (2023)",
"applies_to": [
"broker-dealers dealing with Massachusetts retail customers"
],
"requirement": "A state fiduciary duty of utmost care and loyalty, above the FINRA suitability baseline, upheld by the Supreme Judicial Court.",
"consequence_at_100usd": "Runs in the investor's favour: it is a constraint on the BROKER, protective against gamification and overly permissive options approvals, and creates no filing obligation for the principal. It may however make a Massachusetts firm MORE conservative in granting option privileges than the national convention suggests - which cuts against reaching even Level 2 quickly. Resolved against one source's unsourced negative assertion that Massachusetts imposes no obligation beyond FINRA Rule 2111.",
"status": "settled",
"source_url": "https://www.sec.state.ma.us/divisions/securities/securities-idx.htm",
"retrieved": null
},
{
"authority": "Massachusetts Gaming Commission",
"rule_citation": "M.G.L. c. 23K (Massachusetts Gaming Act, 2011); M.G.L. c. 23N sec. 3 (sports wagering)",
"applies_to": [
"casino and slots gaming",
"sports wagering on athletic contests"
],
"requirement": "Regulates casino gaming and wagering on athletic contests conducted in Massachusetts.",
"consequence_at_100usd": "One source gave c. 23N a different and incompatible identity (a '2016 Fantasy Contest Act' administered by a 'Massachusetts Fantasy Contest Commission') and attributed 2022 sports wagering to a chapter 23O that its own digest flags as nonexistent; both are dropped. Collateral effect: that source's characterization of the Superior Court's c. 23K/c. 23N reasoning is unreliable.",
"status": "settled",
"source_url": "https://malegislature.gov/Laws/GeneralLaws/PartI/TitleII/Chapter23N",
"retrieved": null
},
{
"authority": "Massachusetts Superior Court / Massachusetts Supreme Judicial Court (review pending) / Massachusetts Attorney General",
"rule_citation": "Commonwealth v. KalshiEX LLC, preliminary injunction Jan 2026 (no docket number, division or judge supplied by any source); M.G.L. c. 23K; CEA sec. 2(a)(1)(A) asserted as preempting",
"applies_to": [
"Massachusetts residents",
"CFTC-regulated event contracts"
],
"requirement": "Whether a Massachusetts resident may lawfully trade CFTC-regulated event contracts, and whether the injunction reaches beyond sports contracts to economic, monetary-policy and election contracts.",
"consequence_at_100usd": "THE OPERATIVE GATE, AND THE SINGLE MOST CONSEQUENTIAL DOWNGRADE IN THE MERGED REPORT. Two of three sources characterize the question as unresolved; one stands alone in declaring Kalshi and ForecastEx '100% lawful venues for MA residents'. That claim rests on KalshiEX LLC v. CFTC (D.D.C. No. 1:23-cv-03257, Sept. 12, 2024, 2024 WL 4164694, aff'd D.C. Cir. No. 24-5205, Oct. 2, 2024) - the best-formed case citation in the merge, retained as such - but that decision adjudicated whether the COMMISSION could block election contracts under its own sec. 5c(c)(5)(C) review authority, NOT whether the CEA preempts a state's application of its own gambling law to a state resident, and no authority is offered for the extension. The 2024 federal decisions also predate the 2026 state actions, and the CFTC's amicus filing in the Massachusetts SJC is itself evidence the question was live. Every source describing the injunction describes it as reaching SPORTS contracts and requiring geofencing; whether it reaches non-sports contracts is unresolved by Massachusetts courts. PRACTICAL POSTURE: the conservative reading - that Massachusetts treats event contracts as wagering under c. 23K until a court says otherwise - costs a vehicle; the permissive reading costs potentially a great deal more and rests on a preemption argument no cited authority makes. With USD 100 at stake and legal exposure of unbounded size, the asymmetry strongly favours the conservative reading, not because it is established but because it is the cheap error.",
"status": "unsettled",
"source_url": "https://en.wikipedia.org/wiki/Kalshi",
"retrieved": null
},
{
"authority": "Massachusetts Attorney General",
"rule_citation": "M.G.L. c. 93A (Consumer Protection Act)",
"applies_to": [
"businesses dealing with Massachusetts consumers",
"brokerage gamification practices"
],
"requirement": "Prohibits unfair and deceptive acts and practices; the operative precedent offered is a Robinhood USD 7.5 million settlement (2024) for deceptive gamification.",
"consequence_at_100usd": "Purely personal, self-directed algorithmic execution against a broker's API is UNIMPEDED. Chapter 93A reaches deceptive practices BY a business TOWARD consumers; a principal trading their own USD 100 is on the protected side of the statute, not the regulated side. The exposure flips only if outputs are shared for compensation. The office has taken no public position on non-sports event contracts, and the conservative inference is that it treats all event contracts as gaming under c. 23K until a court rules otherwise. (A c. 12 secs. 4L-5 citation for the office's general authority mixes lettered and numbered sections with no pincite and is downgraded.)",
"status": "settled",
"source_url": "https://www.mass.gov/orgs/office-of-attorney-general-maura-healey",
"retrieved": null
},
{
"authority": "Internal Revenue Service",
"rule_citation": "26 U.S.C. sec. 1222(1) (short-term definition); sec. 1(h) (rate schedule); sec. 1091(a) (wash sale); sec. 1256(a)(1) and (a)(3) (60/40 mark-to-market); sec. 1256(f)(5) (wash-sale exemption); sec. 165(d) as amended by Pub. L. 119-21 sec. 70114(a) (wagering losses)",
"applies_to": [
"all realized gains within the 90-day window"
],
"requirement": "Gains on capital assets held one year or less are short-term and taxed at ordinary rates of 10%-37%. Wash-sale losses are disallowed on stock or securities where substantially identical property is acquired within the 30-day window on either side. Section 1256 contracts are marked to market at December 31 with a 60/40 long/short split regardless of holding period and are exempt from the wash-sale rule. Wagering losses are deductible only to the extent of wagering gains, further limited to 90% of losses.",
"consequence_at_100usd": "WITHIN A 90-DAY EXPERIMENT NO POSITION CAN REACH LONG-TERM TREATMENT - a point worth stating because one source applied 15% long-term rates to its equity baseline, which is impossible on this horizon and inflates the apparent tax advantage of equities. The effective date of the 90% wagering-loss limitation is stated two incompatible ways within one source and must be verified against the enacted text before any after-tax calculation is relied on. Specific subsection letters within sec. 1256(g) are mutually inconsistent across that source's own adjacent table rows and are not asserted; the concepts are settled, the letters are not. FEDERAL CHARACTERIZATION OF PREDICTION-MARKET PROCEEDS IS UNSETTLED - no IRS Notice, Revenue Ruling, Private Letter Ruling or regulation has resolved it despite these contracts existing since 2021, and both sources that address it agree. The decisive practical observation: a retail trader CANNOT INFLUENCE WHICH FORM THE VENUE ISSUES, so the characterization is operationally in the venue's hands.",
"status": "unsettled",
"source_url": "https://www.law.cornell.edu/uscode/text/26/1256",
"retrieved": "2026-08-01"
},
{
"authority": "Massachusetts Department of Revenue",
"rule_citation": "M.G.L. c. 62 sec. 1 (conformity); sec. 4 (gains); sec. 3(B)(a)(13) (gambling winnings at 5.0%); sec. 3(B)(a)(18) and TIR 15-14 (gambling-loss disallowance); Chapter 50 of the Acts of 2023 and TIR 24-4 (short-term rate reduced from 12.0% to 8.5%)",
"applies_to": [
"Massachusetts residents",
"all realized gains"
],
"requirement": "Massachusetts short-term capital gains are taxed at 8.5%; gambling winnings at 5.0%; and Massachusetts DISALLOWS gambling-loss deductions for wagering not licensed by Massachusetts.",
"consequence_at_100usd": "Under a wagering characterization the loss disallowance means the 5.0% state tax applies to GROSS WINNINGS with no offset - a trader can lose money on the year in aggregate and still owe Massachusetts tax on every winning contract. That is a change in the TAX BASE, not a rate difference, and it interacts with the unsettled federal characterization to produce genuine downside asymmetry. Quantified effect: the required break-even win rate rises from 50.00% under capital-gains/sec. 1256 treatment to 51.66% if federal losses are itemized, or 57.80% if they are not (the arithmetic checks internally but the derivation is not shown and the parameters are assumptions). RESOLVED AT [T3] WITH A STANDING VERIFICATION FLAG: TIR 15-14 does not appear in the asserting report's own bibliography, and that report's JSON sources the same finding to a CPA firm's marketing blog. The competing 5% Massachusetts capital-gains rate offered by the other source comes from a passage that source itself flags as unverified and which contains three irreconcilable rate figures; all of that source's after-tax figures therefore understate the Massachusetts component and are not propagated. No MA DOR Technical Information Release addressing event contracts was located.",
"status": "contested",
"source_url": "https://www.mass.gov/technical-information-release/tir-24-4",
"retrieved": null
},
{
"authority": "SEC / Massachusetts Securities Division (Investment Advisers Act)",
"rule_citation": "Investment Advisers Act sec. 202(a)(11), 15 U.S.C. sec. 80b-2(a)(11); registration under 15 U.S.C. sec. 80b-3; M.G.L. c. 110A sec. 201 (Massachusetts RIA registration); publisher's exclusion at sec. 202(a)(11)(D)",
"applies_to": [
"persons who, for compensation, engage in the business of advising others on securities"
],
"requirement": "Four elements must ALL be present: engagement in the business (regularity, not episodic); for compensation (any economic benefit, including indirect, performance-based or revenue-shared); advising OTHERS; on securities (state law often extends further, to commodities and digital assets).",
"consequence_at_100usd": "NOTHING IN THIS EXPERIMENT TRIGGERS A REGISTRATION OBLIGATION so long as the system's outputs stay private and uncompensated - a principal advising themselves satisfies neither the compensation nor the advising-others element. Both sources that address it agree, making this the most reliable conclusion in the regulatory analysis. THE BOUNDARY MOVES THE MOMENT any of three things happen: outputs are shared with anyone (Substack, X, Discord, Telegram, YouTube); compensation of any kind is received (subscription fees, advertising, affiliate or referral revenue from a broker or exchange referral programme, platform-shared revenue, tips); or recommendations are tailored to a specific recipient. At that point federal registration is likely required and distributing signals for compensation triggers Massachusetts state RIA registration under M.G.L. c. 110A sec. 201. The publisher's exclusion is narrow and depends on compensation flowing from subscription revenue rather than advisory fees - sponsorships and affiliate kickbacks can pierce it - and its JUDICIAL CONSTRUCTION IS UNKNOWN here because the only case authority offered was conceded by its own source as 'not retrieved this session' and was dropped entirely.",
"status": "settled",
"source_url": "https://www.law.cornell.edu/uscode/text/15/80b-2",
"retrieved": "2026-08-01"
}
],
"tax_treatment": [
{
"vehicle": "Equity / ETF, held <= 90 days",
"federal_characterization": "Short-term capital gain; ordinary rates 10%-37% (26 U.S.C. sec. 1222(1); sec. 1(h)). Within a 90-day experiment NO position can reach long-term treatment.",
"federal_forms": [
"Form 1099-B",
"Form 8949",
"Schedule D"
],
"wash_sale_applies": true,
"ma_state_characterization": "Short-term capital gain at 8.5% (M.G.L. c. 62 sec. 4; Chapter 50 of the Acts of 2023; TIR 24-4), reduced from 12.0%",
"ma_loss_deductibility": "Full dollar-for-dollar offset against capital gains",
"certainty": "settled",
"source_url": "https://www.law.cornell.edu/uscode/text/26/1222"
},
{
"vehicle": "Listed equity option (long call / put)",
"federal_characterization": "Capital asset; short-term at this horizon; treatment on exercise per 26 U.S.C. sec. 1234",
"federal_forms": [
"Form 1099-B",
"Form 8949",
"Schedule D"
],
"wash_sale_applies": true,
"ma_state_characterization": "Follows federal; MA capital gain at 8.5%",
"ma_loss_deductibility": "Capital loss against capital gains",
"certainty": "settled",
"source_url": "https://www.law.cornell.edu/uscode/text/26/1234"
},
{
"vehicle": "Broad-based index option / nonequity option",
"federal_characterization": "Section 1256 contract; 60% long-term / 40% short-term regardless of holding period (26 U.S.C. sec. 1256(a)(1), sec. 1256(a)(3)); marked to market at December 31",
"federal_forms": [
"Form 6781",
"Schedule D",
"Form 1099-B Boxes 8-11"
],
"wash_sale_applies": false,
"ma_state_characterization": "MA capital gain; 8.5% on the short-term portion",
"ma_loss_deductibility": "Full mark-to-market offset",
"certainty": "settled",
"source_url": "https://www.law.cornell.edu/uscode/text/26/1256"
},
{
"vehicle": "Regulated futures contract (CME futures; KalshiEX BTCPERP)",
"federal_characterization": "Section 1256 contract; 60/40; mark-to-market at December 31. Whether the 90-day window STRADDLES December 31 determines whether mark-to-market bites at all - a straddling open position generates a December 31 recognition event on an unrealized gain with no cash to pay it from. No source states the experiment's start date.",
"federal_forms": [
"Form 6781",
"Schedule D",
"Form 1099-B Boxes 8-11"
],
"wash_sale_applies": false,
"ma_state_characterization": "MA capital gain under conformity at 8.5% short-term portion",
"ma_loss_deductibility": "Capital loss against capital gains",
"certainty": "settled",
"source_url": "https://www.law.cornell.edu/uscode/text/26/1256"
},
{
"vehicle": "Event contract (Kalshi / Polymarket / ForecastEx) - Argument A: Section 1256 regulated futures contract",
"federal_characterization": "Regulated futures contract listed on a CFTC-designated contract market, on a qualified board or exchange, cash-settled at maturity, yielding 60/40 treatment. WEAKNESS: with the CEA sec. 5c(c)(4) citation dropped (that provision concerns self-certification and Commission stay procedures, not mark-to-market), Argument A's mark-to-market premise is unsupported by any retained authority, and CFTC classification is not binding for IRS purposes.",
"federal_forms": [
"Form 6781",
"Form 1099-B Boxes 8-11"
],
"wash_sale_applies": false,
"ma_state_characterization": "MA capital gain under conformity, 8.5% short-term portion",
"ma_loss_deductibility": "Capital loss against capital gains",
"certainty": "unsettled",
"source_url": "https://www.law.cornell.edu/uscode/text/26/1256"
},
{
"vehicle": "Event contract (Kalshi / Polymarket / ForecastEx) - Argument B: wagering transaction",
"federal_characterization": "Ordinary income under 26 U.S.C. sec. 61; losses limited to the extent of wagering gains under sec. 165(d), further limited to 90% of losses by Pub. L. 119-21 sec. 70114(a) - EFFECTIVE DATE UNRESOLVED, stated two incompatible ways within one source. The only authorities offered (Rev. Rul. 54-339 with no bulletin reference or holding, and the 1954 Code sec. 4421 wagering EXCISE-tax definition imported into an income-tax loss provision with no stated bridge) are both [T6].",
"federal_forms": [],
"wash_sale_applies": false,
"ma_state_characterization": "MA ordinary income; gambling winnings at 5.0% (M.G.L. c. 62 sec. 3(B)(a)(13))",
"ma_loss_deductibility": "LOSSES NOT DEDUCTIBLE. Massachusetts disallows gambling-loss deductions for wagering not licensed by Massachusetts (M.G.L. c. 62 sec. 3(B)(a)(18); TIR 15-14), so the 5.0% state tax applies to GROSS WINNINGS with no offset. This is the federal/Massachusetts divergence, and it means a trader can lose money on the year in aggregate and still owe Massachusetts tax on every winning contract - a change in the tax base, not a rate difference. Carried at [T3] with a standing instruction to verify against primary text: TIR 15-14 does not appear in the asserting report's own bibliography, and that report's JSON sources the same finding to a CPA firm's marketing blog [T4]. NOTE ALSO: no source establishes the correct reporting form for this branch - one asserts 1099-MISC Box 3 with no authority while conceding there is 'no specific slot', the other says '1099-B or 1099-MISC', and Form W-2G (the actual gambling-winnings reporting mechanism) is named by neither.",
"certainty": "unsettled",
"source_url": "https://www.law.cornell.edu/uscode/text/26/165"
},
{
"vehicle": "Event contract (Kalshi / Polymarket / ForecastEx) - Argument C: open transaction",
"federal_characterization": "Named as a third possibility in one source's machine-readable appendix and developed by no source",
"federal_forms": [],
"wash_sale_applies": false,
"ma_state_characterization": "",
"ma_loss_deductibility": "",
"certainty": "unknown",
"source_url": null
},
{
"vehicle": "Spot cryptocurrency",
"federal_characterization": "Property / capital asset per IRS Notice 2014-21, with capital gain or loss on disposition; short-term at this horizon",
"federal_forms": [
"Form 1099-DA (brokers in scope, 2025 transactions forward)",
"Form 8949",
"Schedule D"
],
"wash_sale_applies": false,
"ma_state_characterization": "Follows federal; MA capital gain at 8.5% short-term",
"ma_loss_deductibility": "Capital loss against capital gains under conformity",
"certainty": "settled",
"source_url": "https://www.irs.gov/pub/irs-drop/n-14-21.pdf"
},
{
"vehicle": "OTC retail FX / CFD",
"federal_characterization": "26 U.S.C. sec. 988 ordinary by default; election available under sec. 988(a)(1)(B)",
"federal_forms": [],
"wash_sale_applies": false,
"ma_state_characterization": "MA ordinary income under conformity",
"ma_loss_deductibility": "Limited",
"certainty": "settled",
"source_url": "https://www.law.cornell.edu/uscode/text/26/988"
},
{
"vehicle": "Short sale of stock",
"federal_characterization": "Short-term capital gain or loss (26 U.S.C. sec. 1233)",
"federal_forms": [
"Form 1099-B",
"Form 8949",
"Schedule D"
],
"wash_sale_applies": true,
"ma_state_characterization": "Follows federal",
"ma_loss_deductibility": "Capital loss against capital gains",
"certainty": "settled",
"source_url": "https://www.law.cornell.edu/uscode/text/26/1233"
}
],
"backtest_pitfalls": [
{
"pitfall": "Lookahead bias - using information not available at the decision timestamp",
"detection_method": "Interrogate every input: does the dataset at timestamp t contain only information published by t? Run the backtest twice, once with strictly lagged inputs and once with naively aligned inputs, and compare. Mechanically: check whether the feature matrix X_t contains bar-t close/high/low or unannounced fundamental filings. Four retail manifestations: retroactively split-adjusted close prices applied to pre-split decision timestamps; 'most-recent' fundamental values since restated; index membership rebalanced after the backtest window; earnings surprises computed against a consensus not yet aggregated at the decision timestamp.",
"mitigation": "Shift the feature matrix by at least one lag (X_{t-1} -> R_t) and index SEC data by FILING ACCEPTANCE TIMESTAMP, not period end. Lag conventions: one trading day for prices, one business day for fundamentals, one quarter for fundamental filings to absorb the SEC reporting lag (the flat '45 days' stated by one source is the conservative bound, not the rule - 10-Q deadlines are 40 or 45 days depending on filer status).",
"python_tool": "pandas.merge_asof with strict inequality joins; polars.shift(1); pydantic for schema-level enforcement of as-of columns; pyarrow parquet with explicit as-of columns"
},
{
"pitfall": "Survivorship bias - applying a current-universe ticker list to historical data, silently dropping delisted, acquired, renamed and bankrupt names",
"detection_method": "Re-run on a delisting-aware price file and compare equity curves; report the ratio of surviving to delisted names by year; compare the active constituent list against historical delisting archives as of date t. Usable threshold: IF THE SHARPE CHANGES BY MORE THAN 0.3 when a delisting-return adjustment is applied, the backtest was substantially contaminated.",
"mitigation": "A delisting-adjusted dataset - CRSP/Compustat or Norgate point-in-time constituent archives including delisting returns R_delist. The free substitute is a manual construction from SEC EDGAR filing headers carrying delist_date and effective_date, with every row tagged asof_date. At a USD 100 stake the absolute dollar consequence is trivial; the consequence for the INFERENCE is not - the distortion is large enough to flip a deflated Sharpe from positive to negative.",
"python_tool": "duckdb over a locally built security master with historical ticker mapping"
},
{
"pitfall": "Selection bias - choosing a backtest period, asset universe or parameter range AFTER observing which slice produces a positive result",
"detection_method": "There is no post-hoc detection. The only instrument is pre-commitment: pre-register period, universe and parameter grid before running, and report results on a held-out slice. White's Reality Check is the retrospective correction for a candidate pool of known size, and it requires knowing N honestly.",
"mitigation": "A strict out-of-sample window never re-used once a result has been observed, plus an appendix documenting all rejected parameter sets. The retail-practical substitute for academic pre-registration is a decisions.log recording strategy and parameters BEFORE the backtest runs, with an advance commitment to report every pre-registered strategy including the failures. This is the cheapest high-value control in the entire section and the one most reliably skipped.",
"python_tool": "Version-controlled decisions.log; mlflow run tracking"
},
{
"pitfall": "Data snooping and multiple testing - testing many rules and reporting the best as though it were the only one tested",
"detection_method": "Report the number of independent trials N; compute the Deflated Sharpe Ratio; run White's Reality Check or Hansen's SPA against the candidate pool. Operationally, calculate the trial count N and evaluate the VARIANCE OF THE SHARPE RATIOS ACROSS TRIALS, V[{SR_k}], which is the quantity the DSR actually needs. Critical redefinition: for a 90-day experiment THE RELEVANT N IS NOT THE NUMBER OF RULES BUT THE NUMBER OF INDEPENDENT DECISIONS - at most ~63 for a daily-rebalanced strategy.",
"mitigation": "Family-wise error rate or false-discovery-rate correction on the trial pool: Bonferroni (most conservative), Holm (step-down), Benjamini-Hochberg (controls FDR), Romano-Wolf (bootstrap-based, controls FWER under cross-strategy dependence). Report p-values adjusted for N candidates, never raw. Cap the in-sample trial count at N <= sqrt(T) where T is the count of independent returns - for a 63-day equity backtest with T ~ 30 after the autocorrelation haircut that is N <= 5 candidate strategies. (The sqrt(T) bound is uncited in all sources - [T6] - but the arithmetic is internally consistent.)",
"python_tool": "arch.bootstrap SPA / StepM / MCS (first-class classes, contrary to one source's claim that SPA must be hand-rolled); scipy.stats plus a custom Deflated Sharpe module"
},
{
"pitfall": "Overfitting (parameter and feature) - fit flexibility exceeding the information content of the sample",
"detection_method": "Purged and embargoed k-fold cross-validation; report in-sample Sharpe, out-of-sample Sharpe and the ratio; evaluate the IS-vs-OOS Sharpe gap and the combinatorial-CV rank. Compute the Probability of Backtest Overfitting via CSCV/CPCV: PBO = 0.5 means in-sample selection has no better than coin-flip out-of-sample value.",
"mitigation": "Reduce degrees of freedom and penalize the search: L1/L2 regularization, tree depth <= 3 for tree-based learners, and a PBO calculation. Limit engineered features to O(sqrt(T)) where T is the number of INDEPENDENT returns ([T6] - the bound is a heuristic with no source in any of the three reports and is stated inconsistently within one of them). Report BOTH DSR and PBO: DSR tests whether the best in-sample Sharpe differs from zero after correcting for the search; PBO tests whether the best in-sample strategy is THE SAME STRATEGY that is best out of sample. A backtest can pass DSR and fail PBO.",
"python_tool": "skfolio.model_selection.CombinatorialPurgedCV and WalkForward (maintained, BSD-3, sklearn-API-compatible - the single most useful finding in the Python cluster); scikit-learn BaseCrossValidator for a custom purged splitter. NOTE: mlfinlab is NOT open-source software and is unavailable on PyPI"
},
{
"pitfall": "Regime change and non-stationarity - parameters calibrated on a regime that no longer obtains",
"detection_method": "Test parameter stability across rolling windows; apply Chow or Quandt-Andrews breakpoint tests; compute CAGR, Sharpe and tail risk PER REGIME.",
"mitigation": "Three compatible proposals, all retained: Hidden Markov Model regime gating plus crisis sub-sample stress testing; walk-forward with re-estimation frequency matched to the natural regime length, ensembling across regimes, and a hard acceptance criterion requiring A CONSISTENT SIGN OF EDGE IN AT LEAST TWO OF THREE NON-OVERLAPPING PERIODS; and regime-aware models stress-tested across environments.",
"python_tool": "ruptures (CUSUM, PELT, BinSeg) applied both to the strategy P&L series and to each input feature; statsmodels for Chow and Andrews tests; arch for GARCH-derived regime indicators"
},
{
"pitfall": "Transaction-cost underestimation - assuming zero commissions, zero slippage, zero market impact",
"detection_method": "Re-run with explicit per-trade cost (commissions plus exchange fees plus half-spread plus temporary impact) and compute round-trip cost as a percentage of stake; verify the gross-to-net Sharpe decay; compare mid-price execution against the full bid-ask spread and the exchange taker-fee schedule. At a USD 100 stake this is the trap with the largest DOLLAR consequence, because cost is dominated by spread and fixed fees and the spread is a far larger fraction of a USD 100 trade than of a USD 1M trade.",
"mitigation": "Decompose as Cost = Spread/2 + Slippage + Fees. Two falsifiable acceptance rules: require net Sharpe > 0 under conservative costs, and REPORT SENSITIVITY TO A 2x COST ASSUMPTION - a strategy whose Sharpe goes negative at twice the assumed cost is not robust. A backtest that does not report per-trade cost in basis points, with the cost model stated, is unverified.",
"python_tool": "vectorbt / nautilus-trader execution models with explicit fee schedules"
},
{
"pitfall": "Liquidity and market-impact assumptions invalid at retail scale - assuming execution at historical VWAP when the order is a non-trivial fraction of average daily volume",
"detection_method": "Compute the median and 95th-percentile participation rate against 20-day ADV; estimate impact from the square-root law kappa x sigma x sqrt(Q/ADV) ([T6] - presented with no citation and no kappa value in any source, so the functional form is usable and the calibration is not).",
"mitigation": "Cap participation at <= 1% of bar volume / ADV (resolved value; one source floated a looser 1-5% band elsewhere). Restrict to the most liquid ETF and equity subset, with a universe floor of ADV > USD 1M for a 90-day experiment. Add fractional-share routing penalties as an explicit cost line. HONEST CONCESSION PRESERVED: at USD 100 of capital, market impact is usually NEGLIGIBLE on highly liquid instruments (mega-cap US equities, BTC, ETH, SPY, QQQ, TLT, GLD) and binds only on small-cap equities, small-cap ETFs, altcoins and thin-book event contracts. This is the one trap where retail scale is a genuine advantage - and it is exactly offset by transaction-cost underestimation, where retail scale is a genuine disadvantage.",
"python_tool": "Custom participation-rate checks over volume bars; vectorbt / nautilus-trader slippage models"
},
{
"pitfall": "Point-in-time data failures and restatement contamination - using the most-recently-reported figure for a fundamental that has since been restated",
"detection_method": "Compare every input row against the original SEC filing date; for restated values, attach the original filing date as the as-of and ignore later revisions inside the in-sample period. Worked example: an analyst running a value strategy on 31 December 2008 uses the most recent reported book value per share, which reflects impairments and write-downs not filed until the 2009 10-K, and therefore trades on a 'ghost' value.",
"mitigation": "Use the EARLIEST available EDGAR filing of each value, not the latest. Precise join rule: maintain a vintage table keyed by filed_at, each value carrying filed_at and value columns; join the strategy to the value on decision_date >= filed_at, taking the most recent filed_at <= decision_date. Canonical reference implementation: the Philadelphia Fed Real-Time Data Set.",
"python_tool": "duckdb vintage tables over partitioned parquet; FRED/ALFRED vintage_dates parameter for macro series"
},
{
"pitfall": "Backfill bias in vendor datasets - a vendor retroactively adds new listings, splits or index constituents to the historical bar series stamped with original-event timestamps, when no participant could have traded at that price at that time",
"detection_method": "Compare the current vendor universe against contemporaneous vendor snapshots and identify bars absent from the original release; audit vendor schema release histories. Worked example: a vendor adds a stock in 2020 and backfills its price history to 2010 - the history looks complete in 2026, but a researcher who downloaded the same feed in 2012 would never have seen those bars.",
"mitigation": "Use a vendor that preserves vintage history (FRED/ALFRED for macro, CRSP for equities) or maintain timestamped static archives locally; for equities, build the historical universe from EDGAR and refuse to add an asset before its first SEC filing date. THIS IS THE ONLY TRAP IN THE LIST THAT REQUIRES AN INTERNAL ARCHIVE RATHER THAN A THIRD-PARTY PRODUCT - you cannot buy your way out of it retroactively.",
"python_tool": "Immutable local parquet archives written via pyarrow with retrieval_timestamp; duckdb for vintage comparison"
}
],
"honest_conclusion": {
"best_p_reach_200_in_90_days": 0.03,
"corresponding_p_ruin": 0.6,
"positive_expected_value_exists": false,
"summary": "Across the full surveyed universe, the realistic probability that USD 100 becomes USD 200 within 90 days under the best-supported approach is 1% to 8%, with a central estimate of 3%. The corresponding probability of an experiment-killing loss is 45% to 75%, with a central estimate of 60%. No approach in the surveyed universe carries positive expected value after costs and taxes at USD 100 scale. THREE DEFINITIONAL CAUTIONS TRAVEL WITH THESE NUMBERS AND MUST NOT BE DROPPED. (1) 'RUIN' MEANS TERMINAL WEALTH <= USD 25 (an experiment-killing loss), NOT literal total loss; any downstream text labelling 0.60 'probability of total loss' is wrong. Literal total loss is near zero for the top three ranked strategies, which are unlevered spot positions whose modal bad outcome is a partial loss in the USD 40-85 range, and is 0.55-0.92 for long premium options and 0.70-0.90 for event-contract longshots - the approaches that most reliably destroy the entire stake are the ones that most resemble a lottery ticket. (2) BOTH HEADLINE FIGURES ARE UNIVERSE-WIDE AGGREGATES, not row-level values: the single best-ranked row (spot crypto held outright) centrals at P(reach) 0.05 with its own P(ruin) band of 0.40-0.60. The aggregate central sits below the best row's central for two reasons that run in the same direction - that row's band is the least coherently derived in the ranking (its supporting source states the figure three incompatible ways), and the retail reference-class floor of P(reach) < 1% pulls the universe-wide estimate down. (3) THE PORTION OF THE BAND ABOVE ~1% IS ITS LEAST-SUPPORTED PART. The empirical retail floor is P(reach) < 1%; the 3% central sits above it on the strength of documented edges in the ranked strategies, while the same analysis concludes those edges are consumed by friction and tax. Readers should treat 1-2% as the better-anchored end and 8% as the end that depends most heavily on a single incoherently-derived row: THE TRUE VALUE IS MORE LIKELY TO SIT NEAR THE BOTTOM OF THE PUBLISHED BAND THAN THE TOP. The negative-expected-value finding is the only conclusion on which all three independently commissioned reports agree without qualification, reached from three different literatures, three different vehicle universes and three different modelling approaches; it survives every sensitivity the three reports tested, including total collapse of the top-ranked event-contract row on Massachusetts jurisdictional grounds and either direction of the unresolved IRS characterization question. Three mechanisms produce it and they compound rather than substitute: friction drag of 0.02% to 25% of stake depending on vehicle and trade count, with the documented edges concentrated at the expensive end; a tax wedge under which every realized gain inside 90 days is short-term, requiring a gross-up to roughly 1.31x under IRC sec. 1256 treatment and 1.41x if event contracts are characterized as wagering - so an investor who reaches USD 200 gross has NOT reached USD 200, the net position being worth roughly USD 140-171; and a retail reference-class base rate of P(reach) < 1% with P(ruin) 60-70%. Separately and independently of the probabilities: a 90-day USD 100 deployment CANNOT IN PRINCIPLE demonstrate statistical proof of edge, because reaching t >= 3.0 over 90 trading days requires a daily Sharpe of 3/sqrt(90) = 0.3162, an annualized Sharpe of 5.02 that essentially does not exist in unleveraged retail-accessible asset classes. Detecting a +10 percentage-point win-rate edge at alpha = 0.05 with 80% power requires 158 independent trials; the maximum a 90-day experiment can produce is approximately 63, and after the autocorrelation haircut the effective sample is approximately 3 to 30. The experiment's only defensible deliverable is therefore methodological: run it as an instrumented pilot rather than a trial, pre-register before the first trade, and measure realized calibration, realized slippage against mid-price models, realized point-in-time failures and realized cost basis in basis points - all of which are estimable at n = 63 and transfer to a longer-horizon experiment, whereas the dollar P&L does not."
}
}