The deliverable

A document built to be disputed.

Every figure is traceable to how it was produced, every limitation is written down, and the outcome is stated against a rule that was frozen before the result existed.

What is inside

Open any part to see what it contains. Each summary is designed to be enough on its own; the detail is there for when a figure is disputed.

Twelve parts, in the order they are delivered. The verdict is first on purpose: a report that makes you read eleven sections to find the answer is a report written for its author.

  1. VerdictOne result, its reason, and what happens next.

    The first page answers the question that was asked, in one line, in plain language. It states which of the three results applies, why, and what the next step is. Nothing later in the document contradicts it.

  2. Validation Scope, confirmedThe exact claim, asset, timeframe, costs and benchmark you agreed to.

    The scope you confirmed before anything ran, restated verbatim: the strategy version, the asset, the timeframe, the period, the cost assumptions and the benchmark. It sits directly under the verdict because everything below is only valid inside it.

  3. Executive summaryWhat was found, in the length of a page.

    What the strategy does, what was measured, what held and what did not — written for someone deciding rather than someone auditing. The detail is below; this is the part that survives being forwarded.

  4. Key metricsRecomputed from the frozen inputs, reconciled against what you declared.

    Return, volatility, Sharpe, maximum drawdown and turnover, recomputed from the frozen material rather than copied from your report. Where a recomputed figure differs from the declared one, both are shown and the difference is explained.

  5. Benchmark comparisonLike for like, against the comparison agreed at scoping.

    The strategy against the benchmark you confirmed, over the same period, with the same costs and the same exposure basis. Without a fair comparison an improvement cannot be attributed to the strategy rather than to the market it was in.

  6. Classic Monte CarloHow the same edge could have played out along other plausible paths.

    Alternative paths generated while preserving part of the observed time dependence, so drawdown, time under water and losing streaks can be read as ranges rather than as the one number that happened.

  7. Permuted Monte CarloThe same returns, reordered — how much of the risk was sequence luck.

    The observed returns kept exactly, their order changed. If the risk profile falls apart under reordering, the historical experience depended on a sequence rather than on an edge.

  8. Drawdowns and time under waterNot only how deep, but how long.

    Every drawdown with its depth and its duration. Depth is what gets quoted; duration is what decides whether a strategy is actually held long enough to recover.

  9. FindingsWhat was found, located precisely enough to act on.

    Each finding names what was checked, what was observed and why it matters, with the place in the material where it applies. A finding that cannot be located is an opinion.

  10. LimitationsWhat this validation could not establish.

    Gaps in the data, assumptions that had to be made, parts of the claim outside the agreed scope. A limitation is only acceptable if it is specific and written down, which is why this is a section rather than a sentence of general caution.

  11. RecommendationsWhat would have to change, in priority order.

    Concrete, ordered by how much each one would move the result, and written so you can judge whether meeting them is worth the effort before committing to it.

  12. Next stepWhat to do with this, stated explicitly.

    Proceed to unfunded shadowing, revalidate after specific changes, or supply the missing material. The report never ends on an ambiguity about what happens next.

How a result is presented

The same four-pillar structure carries every result. This one is illustrative and deliberately imperfect — a validation that cannot come back inconclusive is not a validation.

Illustrative example — not a client result

Result

Insufficient evidence

Nothing clearly failed, but the material provided cannot settle the question either way. That is the absence of a result, not a negative one, and the two are never reported as if they were the same.

Why
A required category could not be settled either way.
Next step
Supply the missing material so the question can be settled.

How the result was argued

  • CausalPassedEvidence items: 3
  • StatisticalInconclusiveEvidence items: 1
  • EconomicInconclusiveEvidence items: 1
  • RobustnessPendingEvidence items: 0

“Ready for shadowing” is never authorisation to trade real capital. It means the strategy has earned the right to be watched, unfunded, and nothing beyond that.

Validate a strategy — €59Download the sample report

Sample report

I cannot show you a client's report. So I applied the process to a strategy of my own, published openly, and the verdict was that it does not work. Every figure below comes from that validation, and anyone can clone the repository and check it.

VALIDATION REPORT

EWMA63-CURRENT200

A REAL VALIDATION ON A PUBLIC REPOSITORY — ANYONE CAN CHECK IT

Executive summary

The strategy cuts exposure when it detects market stress and, while cutting, reduces high-beta names more. The comparison that decides the question is not buy-and-hold: it is a portfolio that cuts exactly the same exposure without looking at beta. Only then can a difference be attributed to the layer under test.

Strategy under test

CAGR
22.77%
Sharpe
1.343
Max drawdown
-21.82%

Exposure-matched control

CAGR
22.70%
Sharpe
1.324
Max drawdown
-23.46%

Contribution of the layer

Annual active return
+0.02%
Information ratio
0.025
95% bootstrap CI
[-0.31%, +0.38%]

By method

  • CausalPASS WITH LIMITATIONS
  • StatisticalPASS WITH LIMITATIONS
  • EconomicFAIL
  • RobustnessFAIL

Gates

GateWhat it checksThresholdObservedResult
Absolute evidenceWhether the strategy beats the market with significanceHAC t >= 1.963.518PASS
Matched controlWhether the layer adds anything over the exposure-matched controlHAC t >= 1.960.118FAIL
Cost stressWhether that contribution survives a 100 bps trading cost>= 0% at 100 bps-0.043%FAIL
Chronological stabilityWhether at least two of three subperiods are positive>= 66.67%33.3%FAIL

Traceability

Matched control: FAIL leads to Economic: FAIL leads to Revision required for real failure

Limitations

  • The parameters were chosen after seeing the history they are tested on, so no untouched forward period exists.
  • The 200-name universe is companies listed today, which inflates the buy-and-hold reference by construction.
  • The price history contains corporate-action distortions that are documented but not traced to source.
  • The simulation does not model cash return, market impact, order minimums or capacity.

Conditions for revalidation

  1. Freeze the current specification and collect a genuinely prospective period, retuning nothing during it.
  2. Rebuild the universe with issuer identifiers and point-in-time constituents.
  3. Evaluate the exposure gate on its own, since it produces almost all of the improvement.
  4. Add an execution model with the portfolio's real costs and capacity limits.
DeploymentINCONCLUSIVE

INCONCLUSIVE: without a prospective period and a capacity-linked execution model, nothing can be concluded about deployment. It is reported separately from the formal verdict.

Formal verdict

Revision required

Reason: a real economic failure — the layer adds no measurable value over the exposure-matched control

Formal result recorded in the signed report: NO-GO

The full report

This summary is a view of the page. The document a client receives is longer: it explains what the strategy does, what it is compared against and why, and what falls outside the scope. Download it and judge it yourself.

Download the report (PDF)3 pages · PDF

View the repository on GitHub

And an evidence annex

Alongside the report comes an annex holding the series, metrics and figures behind every number: equity curve, drawdowns with their duration, rolling metrics, return distribution, and integrity fingerprints for the input data. The report is what you read to decide; the annex is what holds up when someone disputes a figure.

Extract from the evidence annex that accompanies the report.
Extract from the evidence annex that accompanies the report.

Have something that needs an independent check?

Describe the strategy, the backtest or the model. The scope and the price are confirmed with you before any work starts.

Nothing runs, and nothing is charged, before you approve the scope.

The information you submit will be used to assess your request. Read the privacy policy.