The process

Four steps to a decision, and everything underneath them.

A validation is only worth something if its rules were set before anyone knew the answer. Here are the four steps you buy, the twenty they contain, and the seven contractual stages they run on.

How a validation works

Four steps, from the material you already have to a decision you can act on.

  1. Define

    Send code, Pine Script, a trade list or written rules. Strateva turns the material into an exact scope, and you confirm it.

  2. Execute

    We freeze the rules, the data, the costs and the benchmark. The strategy runs in an isolated environment.

  3. Verify

    The engine reproduces the metrics and examines robustness, time dependence, costs, benchmark and alternative paths.

  4. Decide

    You receive the evidence, the limitations, and a result that tells you what the next step is.

Every validation ends in one of three results

  • Ready for shadowing
  • Revision required
  • Insufficient evidence

AI prepares the strategy. Strateva’s closed engine validates it. Where AI stops

Three evidence levels

Know what kind of evidence you are buying.

Every report names the highest level actually reached. Higher levels require the protocol and time they claim; they are never inferred from a strong backtest.

  1. Included now

    1. Reproduced history

    The frozen strategy is replayed on the confirmed historical data, with declared costs, a benchmark when valid, robustness simulations and a reproducible evidence chain. This is historical evidence, not out-of-sample proof.

  2. When eligible

    2. Walk-forward OOS

    Chronological folds test decisions on later historical observations. Rolling WFA evaluates one frozen candidate; nested WFA is used only when multiple candidates were frozen before selection. It is stronger historical evidence, but still historical.

  3. Requires future data

    3. Prospective holdout

    The strategy, costs, benchmark, horizon and decision rule are sealed before new observations exist. Results remain hidden until the horizon completes. It cannot be delivered immediately or reconstructed from old data.

These are evidence levels within one service, not performance ratings. A result may be unfavourable or not evaluable at any level.

How it is done

Three things almost nobody does, and they are the difference between an opinion and a check.

  • Bias scanner

    Reads your code looking for the future where it should not be.

    An analyser walks the structure of your code, not its text. That matters: the patterns by which the future leaks into the past are almost never written the same way twice, and a text search misses them. It is built assuming nobody writes them obviously — it resolves aliases, computed expressions, and variants that read differently but do the same thing.

    It is a filter, not a certificate. A clean result lowers the risk; it does not prove there is no bias.

  • Rolling or nested walk-forward

    Chronological OOS is used only when the strategy and data support a valid protocol.

    One frozen candidate can use rolling walk-forward without selection. Nested walk-forward requires at least two pre-frozen candidates: inner folds select and outer folds evaluate, with explicit purge, embargo and execution lag.

    Not every validation qualifies. It requires enough observations and does not turn historical OOS into a prospective holdout.

  • Sealed prospective holdout

    The criterion is frozen before the first observation exists.

    The exact future calendar, benchmark, cost model and decision rule are hash-sealed before observations arrive. Metrics remain hidden until the complete frozen horizon exists.

    Time must pass, so it is not part of an immediate historical report. The current internal SHA-256 chain detects ordinary rewrites but is not an independent public timestamp.

Plausible alternative paths

We generate possible alternative paths while keeping part of the time dependence observed in the original series. That makes it possible to estimate how drawdown, time under water, losing streaks and risk of ruin could vary, instead of reading one historical path as if it were the only one available.

The report keeps the decision useful: P5, P50 and P95 are shown instead of a dense catalogue of deciles and percentiles.

How the resampling works

The resampling preserves local structure rather than shuffling observations independently, because independent shuffling destroys exactly the clustering that makes drawdowns deep.

The same result, in a different order

We keep the observed returns exactly as they are and change the order in which they arrive. That shows how much of the risk depended on history having happened in a particularly favourable sequence.

The historical return can stay the same while the experience of risk changes completely.

What this does and does not measure

Permuting the same returns without replacement leaves the terminal result unchanged — the total return is identical in every permutation, by construction. So this is not a distribution of final profit, and it is not presented as one. What it measures is everything that depends on the order:

  • Maximum drawdown
  • Drawdown duration
  • Time under water
  • Longest losing streak
  • Risk of ruin

When many variants were tried

Deflated Sharpe Ratio (DSR) adjusts the selected Sharpe for multiple testing and non-normal returns. PBO/CSCV repeatedly selects the best variant in one half of chronological blocks and checks its relative rank in the other half, estimating how often selection overfits.

This is valid only with the complete, synchronised return matrix for every tried variant and an identified selected variant. One winning strategy, reconstructed alternatives or an incomplete trial history produces NOT EVALUABLE — never an invented reassurance.

The author declares whether the submitted family is complete. Undisclosed trials cannot be detected, and DSR/PBO does not turn historical evidence into a prospective holdout.

How the verdict is built

The result is derived from structured evidence and deterministic rules fixed before execution. It is not a tally of passes: each pillar carries its own conclusion and its own evidence, and the report says which one decided the result. A strategy can clear every statistical check and still fail on economics, and averaging those into one score would hide precisely the thing worth knowing.

Shadowing is not authorisation to trade real capital, and no result guarantees future profitability.

Every engagement ends in one of four outcomes

  • Ready for shadowing
  • Revision required
  • Insufficient evidence
  • Review pending
  • Ready for shadowing: no reason was found to discard the strategy once the benchmark, the costs and its stability over time had been checked. It means “run it in parallel, unfunded, first”. It is never authorisation to go to production or to trade real capital.
  • That outcome can still carry limitations. They are accepted only if they are specific, written into the report, and do not prevent the parallel-running phase.
  • Revision required: the strategy did not pass one of the mandatory checks. The report says which one, what it measured, and what would have to change.
  • Insufficient evidence: nothing clearly failed, but the material does not support a conclusion. That is not the same result as the one above, and you are always told which of the two you got.
  • While automated evidence reconciliation is incomplete, nothing is published as an outcome. Whether the strategy is ready for production is a separate question and is never implied by this validation.

When a strategy does not come through there are two different reasons, and you are always told which one you got: a specific check failed, or the available evidence could not settle the question. “Revision required” and “insufficient evidence” are not the same result, and reporting them as if they were is how a validation stops being useful.

From your material to a verdict

Six steps, and one of them is you. Nothing reaches the engine until you have confirmed what it is about to process.

  1. You

    Submit your material

    Code, Pine Script, a trade list, or rules written in prose.

  2. Claude or ChatGPT, via MCP

    Prepare the material and draft a scope

    Interprets the material and turns it into a draft Validation Scope. It may structure client-provided rules and numerical inputs for review, but it does not calculate metrics, simulations or the verdict.

  3. You

    Approve before anything runs

    Nothing runs before this. You approve the scope, not the AI.

  4. Strateva’s closed engine

    Process the frozen inputs

    The rules, data, costs and benchmark are frozen and run in an isolated environment.

  5. Strateva’s closed engine

    Evidence and deterministic verdict

    Metrics, simulations and a result derived from rules fixed before execution.

  6. Strateva report pipeline

    Final report

    The versioned renderer produces the PDF from verified evidence and the scope it is valid inside.

Where AI stops

Strateva’s quantitative engine is proprietary, closed, and separate from Claude, ChatGPT or any other language model. AI works only as an interaction and preparation layer — it never becomes part of the engine that produces a result.

What it can do

  • Collect the request through MCP.
  • Interpret code, Pine Script or rules written in prose.
  • Detect missing or ambiguous information.
  • Structure the rules and numerical inputs the client provides — capital, costs, slippage, exposure, position sizing, dates and parameters — for review.
  • Turn the material into a structured specification.
  • Prepare code or a reviewable runner when one is needed.
  • Prepare the inputs so the pipeline can ingest them.
  • Explain the confirmed scope and already-produced results to the client.

What it can never do

  • Access the engine’s internal code.
  • Run inside the quantitative kernel.
  • Modify the methodology.
  • Select which metrics get computed to favour a result.
  • Compute the final metrics.
  • Generate the Monte Carlo results.
  • Choose or change the random seeds.
  • Change data, costs, benchmark or criteria after the freeze.
  • Alter the gates.
  • Approve a strategy.
  • Issue, change or overturn the verdict.
  • Invent figures to complete a report.

You confirm the interpretation before anything runs.

Metrics, simulations and verdicts are produced outside the language model by a closed, deterministic validation engine.

Claude or ChatGPT do not call the engine directly and receive no access to its internals. MCP exposes only controlled operations and public contracts — never the kernel itself.

What is checked before delivery

Closed contracts verify the evidence hashes, the relationship between scope and execution, metric reconciliation, the deterministic conclusion and the presence of explicit limitations. Routine manual analyst review is not part of the sponsored product.

Proprietary implementation. Inspectable evidence. Deterministic verdict.

The implementation is proprietary. The engagement is not a black box.

Strateva does not expose its internal engine code. It exposes what matters for reviewing the result:

  • The Validation Scope, confirmed by you.
  • The data, costs, benchmark and criteria used, declared.
  • The process that was executed, documented.
  • The delivered metrics, linked to evidence.
  • The simulations, with their parameters and seeds recorded.
  • The findings and limitations, explained in the report.
  • The verdict, following deterministic rules.
  • What was validated — reviewable even without the engine’s code.

AI and the engine, in the questions we get asked

  1. Does AI validate my strategy?

    No. Claude or ChatGPT may help structure the submitted rules and prepare the material for ingestion. Validation takes place outside the language model. Strateva’s closed quantitative engine calculates the metrics and simulations, and deterministic rules produce the verdict.

  2. Can Claude or ChatGPT access the engine?

    No. They interact through controlled MCP operations and public contracts. They cannot inspect the engine source code, change its methodology, modify frozen inputs or influence the verdict.

  3. Is the engine open source?

    No, the implementation is proprietary and closed. That is a separate question from whether a result can be checked: it can — the confirmed scope, the declared assumptions, the evidence and the verdict rules are documented in the report, without publishing the engine’s code.

  4. If the engine is closed, how can I review the result?

    The implementation is proprietary, but the engagement is not a black box. Your confirmed scope, data and execution assumptions, benchmark, criteria, evidence, limitations and decision trail are documented in the final report.

The twenty steps, in full

The four steps above are what a validation is. These twenty are what it does. They are published so the result can be argued with, not because you need to read them before connecting.

  1. Material received

    Your code, Pine Script, trade list or written rules arrive as they are. Nothing has to be tidied first: the state it comes in is what tells us how much reconstruction the work needs.

  2. Assisted interpretation

    Rules written in prose are read and turned into a draft specification. This is the one step where a model is involved, and it produces a proposal, never a number.

  3. Validation Scope drafted

    The claim under test, the asset, the timeframe, the period, the costs and the benchmark are written down as one exact document.

  4. You confirm it

    Nothing runs until you agree that the specification is your strategy. This is the last moment at which any of it can be chosen.

  5. Data prepared and frozen

    The series is assembled, checked and fixed. From here the inputs cannot move, which is what makes the answer repeatable.

  6. Strategy normalised

    The strategy is expressed in a form the engine can execute without altering what it does.

  7. Runner contract checked

    Before execution, mechanical guards verify the adapter boundary, prohibited operations, hashes and expected output. This is an automated gate, not routine human code review.

  8. Isolated execution

    The run happens in a clean environment with no access to anything outside the frozen inputs, so what is measured is the strategy and not its surroundings.

  9. Strategy Trace produced

    A record of what the strategy did, decision by decision, rather than only what it returned.

  10. Strategy Trace verified

    The trace is checked against the specification: the positions taken are the positions the agreed rules imply.

  11. Quantitative kernel

    The metrics are recomputed from the run and reconciled against what was declared before it.

  12. Classic Monte Carlo

    Alternative plausible paths, to read drawdown and time under water as ranges rather than as the one history that happened.

  13. Permuted Monte Carlo

    The same returns in a different order, to separate an edge from a fortunate sequence.

  14. Evidence versioned

    Every figure is stored with what produced it, so a number can be traced rather than trusted.

  15. Independent verification

    The evidence is checked against the criteria by a path that does not reuse the code that produced it.

  16. Criteria evaluated

    The frozen decision rule is applied. The criteria were fixed before execution, so the result is not chosen once the answer is visible.

  17. Technical result

    The internal outcome, in the engine’s own vocabulary and at its own level of detail.

  18. Public projection

    That result is projected into the published contract: four categories, one status, one reason, one next step.

  19. Automated evidence reconciliation

    Versioned contracts reconcile hashes, metrics, evidence and the frozen decision rule before a result can be published.

  20. Final delivery

    The report, its annex and the result reach you together, with the scope they are valid inside.

Steps are described in public terms. Thresholds, internal identifiers, commands and infrastructure stay where they belong. The sponsored workflow is automated; client confirmation authorises the already-frozen scope.

The seven stages

The order is fixed. No stage is skipped, and once the inputs are frozen no earlier stage can be reopened without starting again.

  1. Submit your strategy

    You

    You describe what the strategy does and send whatever you have: a return series, a notebook, a repository, or the rules written out. Nothing has to be tidy yet — what it is in is what tells us how much reconstruction the work needs.

    Produces: your material recorded as submitted

  2. Confirm the scope

    Together

    We agree in writing on the claim being tested, the benchmark it is tested against, the period, the cost assumptions and what would count as a failure. This is the last moment at which any of it can be chosen.

    Produces: a Validation Scope you have confirmed

  3. Freeze the inputs

    Strateva

    The agreed material, the data and the decision rule are fixed and fingerprinted. From here on, the question is settled and only the answer is unknown.

    Produces: rules, data, costs and benchmark fixed and fingerprinted

  4. Run in isolation

    Strateva

    The strategy runs in a clean, isolated environment with no access to anything outside the frozen inputs, so what is measured is the strategy and not its surroundings.

    Produces: a completed run in an isolated environment

  5. Validate quantitatively

    Strateva

    The results are put through the checks agreed at scoping: the causal structure, the statistical evidence, the economics after costs, and the stability of all of it under stress.

    Produces: the measured evidence behind every check

  6. Review the evidence

    Strateva

    Every check is reconciled mechanically against the evidence behind it. A number that cannot be traced back to how it was produced does not make it into the report.

    Produces: that evidence reconciled by versioned contracts

  7. Receive verdict and report

    Strateva

    You receive one outcome, the reason for it, what would have to change, and the evidence it rests on. The outcome is computed against the rule frozen at stage three, never one chosen once the result was visible.

    Produces: the report and the result

Who does what

What you provide

  • The strategy itself, in whatever form it exists: returns, code, a notebook or the written rules.
  • What you believe it does, and what result you expect it to hold up to.
  • The data, or enough about it that an equivalent set can be assembled.
  • Anything already known to be wrong with it. Declared limitations cost nothing; ones discovered during validation can send the scope back and cost another validation.

What you confirm before anything runs

  • The exact claim under test, written in one sentence.
  • The benchmark it will be compared against, and why that one is the fair comparison.
  • The period, the universe and the cost assumptions.
  • What would count as a failure — agreed while the answer is still unknown.
  • The scope and the price. Neither changes afterwards without your approval.

What Strateva does

  • Freezes and fingerprints the inputs, so the question cannot drift towards the answer.
  • Reproduces the result in an isolated environment, without quietly altering the original claim.
  • Challenges it on all four pillars, not only the one that looks weakest.
  • Reconciles every published number against versioned evidence and hashes.
  • States the outcome against the frozen rule, and says which of the four it is.

What you receive

  • One outcome, in plain language, with the reason it landed there.
  • The four pillars, each argued separately, so a weakness is located rather than averaged away.
  • The limitations and the risks that were not resolved.
  • The conditions under which the result would be worth revalidating.
  • The evidence annex the figures come from.

Why the inputs are frozen

Almost every backtest that fails in production failed the same way: the rule was adjusted, in good faith, after the result was visible. A threshold moved, a period was trimmed, a benchmark was swapped for a kinder one. Each change is defensible on its own, and together they guarantee a good-looking answer.

Freezing the inputs makes that impossible rather than discouraged. The claim, the data, the costs and the decision rule are fixed and fingerprinted before execution, so the outcome is computed against a rule that existed before anyone knew what it would say.

It is also what makes the result repeatable. Anyone holding the same frozen inputs gets the same answer, which is the difference between a validation and an opinion.

What a validation does not do

  • It does not predict future profitability. It reports what held up under the checks that were agreed, over the period that was examined, and nothing beyond that.
  • It is not investment advice and never a recommendation to allocate capital.
  • “Ready for shadowing” is not authorisation to trade real money. It means the strategy may be run unfunded, in parallel, so its live behaviour can be observed.
  • A clean result lowers risk; it does not prove the absence of a flaw. No finite set of checks can.
  • Whether a strategy is ready for production is a separate question, reported separately, and it does not change the outcome.

Have something that needs an independent check?

Describe the strategy, the backtest or the model through the MCP. You confirm the exact scope before anything runs.

Validation is sponsored. Optional support is separate and never affects the result.