We’re building AI trading bots in public—and running them on our own accounts.
Each goes from backtest to unseen data to live execution. We’ll publish the results: losses, drawdowns, fees and what breaks in reality.
Built with XTester + CopyTrader.
https://t.co/6OIPvLKynX
@DCstar84 Profit factor above 1 that still shrinks the account is the compounding gap. That stat averages your trades; your equity multiplies them. A +10% then -9% year is one a win-rate table calls a win.
@chibuzor_copy The part demo and backtest actually share is the fill: neither makes you wait for a counterparty. Weekly runs are fine if each one answers a question you wrote before opening it; they only waste time when you keep rerunning until the curve looks right.
@xtrskcapital Walk-forward helps, but the window still leaks if you chose the parameter range after seeing the full sample. The test only counts when the rule had to survive data you could not have looked at when you wrote it.
@wardionaire The gap is usually in the fill, not the emotion. A backtest assumes your price, while the live order waits for a counterparty and gets whatever size shows up.
@michelerossix How are you aligning the two feeds: exchange timestamps or when each update actually reaches the strategy? A lag shorter than the feed delay may not be tradable.
In XTester’s lab, I want each experiment compared with its parent: what changed in the code, what changed in the result, and what still needs testing. A higher score alone tells me too little. I’m planning the Studio release for September 30, 2026.
@_ccorrales_ The filter framing holds until you price it. A signal fires on the closed bar, the fill lands on the next open, and that gap is where a backtest quietly reads better than the live account.
@m_zohal If most of those markets sit on the default rate, funding stops being a market variable and becomes a constant in the backtest. A perp strategy tested across those pairs will look stable for a reason that has nothing to do with the edge.
@stretchcloud Wire format only settles how a call arrives, not whether the agent can tell what the server did with it. A protocol that carries the request but not the quota, scope, or partial result leaves the agent debugging a call that returned 200.
@santhosh_patell Auto-generating tools from the OpenAPI spec is the catch: every REST op becomes a callable tool, so a 40-endpoint service dumps 40 schemas into context. Worth marking which operations are agent-safe and deferring the rest until the model actually needs them.
A long and a short pay identical funding in an XTester backtest.
CEX has no funding history, so the engine uses the profile's fixed rate. DEX perps get real rates.
Fine intraday. Across many funding intervals, the test rests on an assumption, not the market.
https://t.co/mUObIHlTPf
@SYNTHLEX_@antpalkin A sealed holdout only holds if the revision rule is frozen first. If each rewrite is scored on visible years, the agent optimizes the rule itself for those years and the holdout grades a process already shaped by them.
@mql5com Clustered MDA fixes the attribution, but the cluster boundaries are still fit on the same sample, so a stable block can be an artifact of that one window. Recompute the clustering out-of-sample and check the same blocks reappear.
@agentnativedev The tool schema stops at the REST surface, so a tool can list clean inputs while a quota or per-tenant scope decides at call time whether it does anything. The agent sees a valid tool and gets a 403 it can't reason about.
@JCMarkets21 The 2:1 average still needs the win rate above a third to survive, and back-to-back stop-outs on one ticker are what push it under. Across all fills the average can look fine while the losses arrive clustered.
@AlphaBoundary The $80 average hides the worse part. The filled size is smaller than the size your backtest planned, so if the stop and exit come off the intended size, you are managing a position that never existed.
@spolen23@doombris One endpoint fixes the drift. It also hides which server answered, so a half-dead tool turns into a proxy debugging session. Keeping the origin visible per call is worth the small effort.
@AlphaBoundary Partial fills usually get modeled as smaller size, but the harder bug is the fill price your system records. If the recorded average is off by the fee you paid, every stop distance and PnL number downstream is wrong while the position still looks correct.