AcademyReference
Reading a result
How to interpret a TextToQuant backtest and report it honestly: what the grade means, what the honesty flags oblige you to say, and the order to say it in.
Written for whoever does the reporting, which is often not the person who ran the backtest. If you are an AI agent relaying a result, this page is also served as Markdown.
The order
Say it in this order.
The headline number is the last thing that should come out of your mouth, not the first.
- 01
Trade count first
Under 30 trades, say so before anything else and say plainly that the sample does not establish an edge. Under 15, the grader has already multiplied the score by 0.6 and the run is a curiosity, not a result.
- 02
Then the honesty flags
Every warning on the run, in plain words. If costs were not modelled, that comes before the return, not after it.
- 03
Then the grade and the badge
The letter and DEPLOY or AVOID, with the pillar that dragged it down. The grade already accounts for sample size, so it is more honest than the raw return.
- 04
Then the numbers
Return against max drawdown, never return alone. A 300% return with a 90% drawdown is not a good outcome.
- 05
Then what would raise confidence
One concrete next test: a parameter sweep, a walk forward, an overfit audit. Say what it would cost.
Sample size
Too few trades is a number, not a feeling.
The grader multiplies the score by this before anything else, so a thin run is already marked down by the engine. Quote the trade count whenever you quote the grade.
| Trades | Score multiplier |
|---|---|
| 0 to 14 | ×0.60 |
| 15 to 29 | ×0.75 |
| 30 to 49 | ×0.90 |
| 50 to 99 | ×0.95 |
| 100 or more | ×1.00 |
The grade
What the letter means.
A composite of four pillars (profitability, risk, consistency, edge), each scored 0 to 100, with the sample size multiplier above applied. 65 or higher carries the badge DEPLOY; below it, AVOID.
Honesty flags
What each one obliges you to say.
Graded without transaction costs / without trading fees
No fee, commission or slippage was modelled, so every trade was free. Our own corpus audit found that switching costs on flipped better than one strategy in five from profitable to unprofitable, and moved the median run down about 10 percentage points.
Say: State that costs were not modelled BEFORE quoting the return, and say the real result would be worse. Do not present a frictionless run as a result.
No losing trades: grade is PROVISIONAL
There is no observed risk unit, so R is estimated from adverse excursion or drawdown. A loss free backtest usually means overfitting, look ahead bias, or too few trades. The score is capped at C.
Say: Say the grade is provisional and why. A strategy that never lost is a warning sign, not a selling point.
R calculated from only N losing trades
Fewer than 10 losses, so expectancy and risk reward rest on a handful of observations.
Say: Quote the number of losing trades alongside any expectancy or R multiple figure.
The return magnitude this grade anchors on is EXTRAPOLATED
The run covers less than a year, so its pace was compounded to a full year the strategy never traded. The annualized figure describes a period that did not happen.
Say: Never quote the annualized number on its own. Give the window length and the actual return over it.
Monte Carlo not run; using equity curve R² as consistency proxy
The consistency pillar is an approximation. The full grade needs Monte Carlo.
Say: Mention that the grade is partial and offer to run the full analysis.
MC ruin > 40% / > 20% / MC median return < 0
Monte Carlo resampling says a meaningful share of orderings end in ruin, or the median ordering loses money. The score is hard capped at D, C+ or C respectively.
Say: Lead with the ruin probability. A strategy whose median resampling loses money is not a strategy that made money.
Metrics
The ones worth naming.
Worked example
The same run, reported two ways.
totalReturn +340%, trades 8, fees_bps 0, stop none, window 5 months, grade D, badge AVOID
Wrong
This strategy returned 340%, which is an excellent result.
Right
This run took only 8 trades, which is far too few to conclude anything, and it modelled no trading costs. The engine graded it D and flagged it AVOID. The headline +340% comes from 8 trades over 5 months with no fees, so treat it as a curiosity rather than an edge. If you want to know whether there is anything here, the next step is to re run it over a longer window with realistic costs.