F1 ยท Publication Volume 30

Holdout Validation

freezing a judgement, introducing independent data, scoring and failure analysis

Learning objectives and boundary

  • Explain freezing a judgement, introducing independent data, scoring and failure analysis as an auditable evidence problem.
  • Separate source observations, controlled derivatives, interpretation and decision state.
  • Apply identity, spatial, semantic, lineage, quality and review checks.
  • Produce a sealed submission manifest, transparent scorecard and immutable failure analysis.

This lesson is part of a general, institution-neutral tutorial and has no relationship to any company or individual. All identifiers, geometries and values in the worked case are explicitly synthetic. The method does not confer professional authority and must not be transferred to a real decision without appropriate sources, competence, law and review.

Core method

Holdout validation asks whether a frozen interpretation generalises to independent evidence. Before reveal, freeze the charter, source snapshot, transformations, hypotheses, target claims, confidence categories, scoring rule and submission digest. Protect the holdout from development paths and record who performed generic custody, reveal and scoring roles without making any individual or organisation part of the tutorial.

Independence is geological and informational, not merely a different filename. Spatial neighbours, later versions, derived summaries and shared source lineages may leak target information. Reveal once, score every prespecified claim including abstentions, and preserve failures. After reveal the data become evaluation history, not a fresh holdout for the next version. Failure analysis separates source gaps, spatial errors, transformation errors, hypothesis errors, calibration errors and review failures.

A sealed validation sequence from frozen claims through independent reveal, scoring and failure analysis
A sealed validation sequence from frozen claims through independent reveal, scoring and failure analysis

Record and evidence model

| Field or object | Operational meaning | |---|---| | submission_id and digest | exact frozen workbench state | | partition_id and independence_basis | protected set and separation rationale | | claim_id, prediction and confidence | prespecified evaluand | | reveal_event and score_rule | controlled access and fixed evaluation | | failure_class and remediation | preserved error and proposed learning |

Every record carries a version, validity interval, source or derivation link, quality state and review state. Missing, unknown, not applicable and failed are separate states; none is silently represented as zero or an empty string.

Controlled workflow

  1. Define geological and informational independence.
  2. Freeze the full evidence state and scoring rule.
  3. Bind the submission digest before reveal.
  4. Reveal the protected data once under recorded custody.
  5. Score every prespecified claim, contradiction and abstention.
  6. Classify failures and reserve fresh evidence for the next version.

Record each step as an activity with declared inputs, parameters, outputs and checks. A rerun creates the same content or a documented difference; it never overwrites the evidence needed to explain the previous state.

Worked synthetic example

SYN-WB freezes three target statements and one abstention. The sealed package contains two synthetic holes positioned by spatial blocks outside the development lineage. Reveal supports one structural prediction, contradicts one alteration-continuity claim and leaves one depth prediction unresolved because recovery failed. The unresolved result is not scored as absence. The failed claim remains in the report, and the used holes are marked evaluation history.

The example is deliberately small enough to inspect row by row. Its names and numbers are fictional teaching values and cannot be used to infer a real place, right, resource, hazard or organisation.

Quality control and failure modes

Run six gates before accepting the lesson artefact: identity resolves the exact objects; spatial control confirms coordinate and extent validity; semantics preserve units, vocabularies and epistemic class; lineage exposes every transformation and dependency; quality records passed, failed and not-run checks; review binds a disposition to the exact content digest. A failed gate is retained as evidence and blocks only the affected use. It is never converted into a favourable value or hidden to simplify the display.

Common failure modes:

  • Choosing the holdout after viewing outcomes
  • Using random rows despite spatial dependence
  • Changing thresholds after reveal
  • Deleting failed predictions from the denominator
  • Reusing revealed evidence as a fresh blind test

Practical exercise

Complete the following tasks against a new copy of the synthetic package:

  1. Design a spatially independent synthetic holdout.
  2. Write a pre-reveal submission manifest.
  3. Score support, contradiction, unknown and abstention.
  4. Classify failures without rewriting the frozen hypothesis.

For every answer, cite the package object IDs, show the failed as well as passed checks, and state which conclusion would change if one assumption were reversed.

Completion artefact and verification

Deliver a sealed submission manifest, transparent scorecard and immutable failure analysis. Include a manifest, exact input identities, transformation or reasoning record, machine-readable validation results, human-readable limitations and the current review state. A second reader must be able to trace one accepted conclusion to its source support, one rejected path to its first failed gate and one unknown to the evidence required to resolve it. Rebuild the artefact from a clean location and compare content digests. Completion is withheld when a required source, parameter, permission or review is missing; no substitute data are invented.

Sources