Building open-source tools for trustworthy quantitative research. Backtesting • causal execution • Python • research infrastructure. Building in public.
Most backtests don't fail because of bad math. They fail because the research pipeline quietly cheats.
I build open-source tools and document real quant research: causal execution, data integrity, strategy robustness, and what breaks in OOS.
Thread ↓
@quantbeckman Would a pre-specified regime rule and untouched walk-forward test change your view, or is the gain over a random walk too small to matter?
@didier_lopes The editable derived widgets are a big step. For quant research, the hard part is preserving lineage: source/version, as-of timestamps, join keys, and transformation history. Will the workspace expose that provenance so results can be reproduced?
One rule I use in quantitative research:
completed-bar signal → execute on a later bar
No future high/low.
No same-close fill after reading that close.
No hidden intrabar assumptions.
If the simulator can’t explain when information became available, I don’t trust the result.
@ml4trading@MavenHQ This is where agent speed needs a hard evaluation harness. If the agent can iterate on the same backtest repeatedly, it can overfit faster than a human. I'd lock data boundaries, costs, and OOS rules before the agent sees any performance feedback.
@pyquantnews The software stack can cost $0. The research stack usually doesn't. Clean point-in-time data, corporate actions, survivorship control, realistic execution, and validation discipline are where many backtests become expensive—or misleading.
@QuantInsti The "miss the best days" result is striking, but it's also ex-post selection. A cleaner test would compare precommitted de-risking rules vs buy-and-hold using the same data, costs, and re-entry logic. That tells us whether avoiding drawdowns actually improves the path.
@quantpedia A useful addition would be to split "Robustness" into components: parameter stability, time-split consistency, cost sensitivity, and regime dependence. A single robustness score is convenient, but it can hide which failure mode is actually driving the result.
A backtest can be mathematically correct and still be impossible to trade.
The first thing I check is causality:
If a signal uses a completed bar’s close, execution cannot happen at that same close.
Small timing assumptions can create very large fake edges.
@quantpedia Interesting way to escape the short history of social-media data. One robustness question: as the feature set changes across eras, does the classifier's decision boundary drift too? Walk-forward re-estimation seems important so "memeness" doesn't quietly become a regime proxy.
@QuantInsti The Germany + spread decomposition is the useful part here: it separates the common global-rate factor from France-specific risk. I'd also check rolling correlations and stress-event windows—full-sample averages can hide state dependence exactly when contagion matters most.
@pyquantnews I’d add one more test: can you explain exactly when information becomes available and when execution is allowed? A simple rule can still hide look-ahead if the timing is vague. Simplicity helps, but explicit causal timing is what makes a backtest believable.
If you work on systematic trading, Python, market data, or open-source research, follow along.
GitHub: https://t.co/LzIzStXJRJ
Contributors and critical feedback are welcome.
Most backtests don't fail because of bad math. They fail because the research pipeline quietly cheats.
I build open-source tools and document real quant research: causal execution, data integrity, strategy robustness, and what breaks in OOS.
Thread ↓
Current open-source projects:
• Backtest Integrity Guard
• Causal Backtest Harness
• Canonical Ledger Schema
• Strategy Stability Report
Built to make backtests easier to audit — and harder to fool.
A small but meaningful milestone for Backtest Integrity Guard: the first community contribution has been merged.
A new contributor picked up a good-first issue, forked the repo, implemented `btguard --version`, passed CI on Python 3.10/3.11/3.12, and the PR is now in main.
3/ Same-bar exits
If both stop and target are touched inside one OHLC bar, the intrabar path is unknown.
The tool flags:
AMBIGUOUS_SAME_BAR_EXIT
unless an explicit policy is declared.
Three backtest failures I now check before trusting an equity curve:
1) same-bar execution after using the close
2) missing OHLCV bars
3) stop + target touched in the same bar
I open-sourced a small Python CLI to fail closed on these cases. Thread ↓
2/ Missing bars
Valid OHLC geometry does not mean the series is complete.
For a declared 5-minute cadence, a jump from 00:00 → 00:15 becomes:
MISSING_BARS
Known session gaps can be allowlisted explicitly.