Thales
methodology & verdicts

Research, honestly

Most quantitative trading research you can find online shows you its winners. This page works the other way: it lists the rules that make results trustworthy, and then every idea we built that failed — because a research process that cannot show you its dead ends is indistinguishable from luck.

The backtest is a falsification engine, not an optimizer.

Historical simulation is only trusted here to KILL ideas. A good backtest makes an idea eligible for live paper trading; it is never treated as a prediction of profit. The strongest force in quantitative research is self-deception, and this rule is the main defense.

Rules are written before results exist.

Pass/fail criteria — including the December go-live gate and each experiment's test battery — are written down and frozen before the data that will judge them arrives. Deciding the rules after seeing the outcome is how every fake track record in this industry gets made.

Tests run on survivorship-free data.

Backtests here include the stocks that later went bankrupt or were delisted — the ones most datasets quietly forget. Including the corpses typically makes results substantially worse, and substantially more honest.

Headline numbers ship with their asterisks.

Example, applied to our own flagship: the momentum backtest's headline is the LUCKIEST of eight calendar start-date variants — across all eight, the average result is roughly zero. We publish that spread alongside the headline, because a number without its luck estimate is advertising, not research.

A negative control runs live.

One sleeve trades daily under a strategy our tests rejected, with a covenant that it can never be promoted. If it quietly wins, our tests are broken. No result from the other sleeves is fully trusted unless the control behaves as expected.

the kill list

Built, tested, dead

Every strategy idea this platform built and tested, with the honest verdict. Publishing failures is the point: a research process that only shows winners is indistinguishable from luck. Verdicts come from pre-registered statistical gates (overfitting probability, out-of-sample tests on survivorship-free data), not from taste. Much of this list was built, tested, and killed by the AI agent that operates the platform, during overnight research sessions run under the same pre-registered rules.

Machine-learning trade filter (random forest meta-labeling)2026-05Killed

Overfitting probability measured at 100% with price-only features — the model memorized the past instead of learning anything. Disabled in config; would only be revisited with fundamentally new input data.

Residual momentum (factor-neutralized ranking)2026-05Killed

Worked mechanically exactly as the literature describes — and cost about 1.7% of annual return in practice. An idea can be real and still not worth its price.

Value sleeve (cheap-stock tilt)2026-05Killed

Added no robustness benefit on the overfitting tests. Removed.

Quality, accruals, asset-growth and low-volatility tilts2026-05Killed

Every classic fundamental tilt was built and tested; none survived the honest evaluation harness. The conclusion — published price+fundamental signals on US large caps are exhausted — is now a standing research rule here.

Graduated index-hedge overlay (scaling into SPY in stress)2026-05Killed

A return drag: about −1% annualized in walk-forward testing, losing in 0 of 5 test windows. Protection that costs more than the crashes it softens.

A faster variant of the flagship's signal2026-05Not shipped

Worsened the overfitting probability in both test modes. The incumbent stayed.

A three-parameter variant of the flagship2026-05Not shipped

Looked better on the headline number, but the verdict flipped depending on how the test data was partitioned — the signature of a fragile result. Investigated, documented, not shipped.

Tranching (staggering the rebalance across the month)Jun 10, 2026Killed

Killed by its own pre-registered gate before launch. Its useful by-product survived: a diagnostic that measures how much of any backtest headline is start-date luck — now run against our own numbers.

Corporate-insider trading signal (SEC Form 4 filings)Jun 10, 2026Killed

A perfect null: after building the full data pipeline, insider filings predicted nothing at all in our universe. The cleanest kind of kill.

Short-horizon mean reversionJul 5, 2026Repurposed

Rejected by its pre-registered test battery — then deliberately deployed live anyway as the negative control, under a covenant that it can never be promoted. If our rejected strategy quietly wins, our tests are broken. It trades every day to keep that check honest.

Micro-cap reversal program (v1)Jul 22, 2026Killed

A month of data engineering (point-in-time small-cap universe, delisted-stock recovery, realistic cost model), then one hash-frozen test battery allowed to run exactly once. It failed: out-of-sample Sharpe −0.30, negative even at half the estimated trading costs. Killed without appeal; any v2 requires a brand-new pre-registration.

Learning-to-rank stock selection (LambdaRank)2026-05Paused

No honest gain over the simple momentum ranking to date. Archived, not deleted — would be revisited only with genuinely new input data.

Congressional-trading data (as a signal source)Jul 25, 2026Killed

Evaluated during a data-sourcing deep-research pass: since the 2012 disclosure law, politicians' disclosed stock picks slightly underperform. A famous signal that died the moment it became famous.

still standing

What survived

The configuration running in production today.

What survived every gate and runs in production today is a single trend-following configuration on large-cap US equities, wrapped in layered risk controls: diversification-aware weighting, conservative sizing, volatility targeting, and a hard drawdown stop. The exact signal and parameters stay in the private repository — what we publish is every output it produces, every day, and the honest scorecard including its caveats.