E7 · Publication Volume 29
Competing Hypotheses and Structured Reasoning
predictions, support, contradiction and missing evidence
*predictions, support, contradiction and missing evidence*
Learning objectives and boundary
This lesson is general and institution-neutral. It uses no real company, individual, property, project or identifiable place. Generic roles describe responsibilities only, and every SYN-AI identifier denotes explicitly synthetic teaching evidence.
- Frame the decision governed by predictions, support, contradiction and missing evidence.
- Separate generated proposals from admissible evidence, deterministic results and accountable judgement.
- Define measurable failure, abstention, escalation and release conditions.
- Produce a reviewed competing-hypothesis matrix with frozen predictions and test plan from synthetic evidence and defend its controls.
Decision and professional boundary
Reasoning is structured around alternatives before evidence is scored. State at least two plausible hypotheses, the scope of each and the observations each predicts. Then classify evidence as supportive, contradictory, non-diagnostic or missing relative to each prediction. Do not ask a model to invent a single geological story first and search for confirming details afterward.
The objective is not to make every hypothesis equally likely. It is to expose which observations would discriminate among them, where dependencies make evidence less independent and what result would cause revision. Generated suggestions may help enumerate alternatives, but hypotheses, predictions and scoring rules are reviewed and frozen before the decisive evidence is revealed.
Core concepts
| Concept | Operational meaning | |---|---| | Hypothesis | A bounded explanatory proposition that makes testable predictions. | | Prediction | An expected observation with location, scale, condition and tolerance. | | Diagnostic evidence | Evidence whose possible outcomes differ meaningfully among hypotheses. | | Contradiction | An admissible observation inconsistent with a declared prediction under its scope. | | Unknown | A required test or observation that has not been obtained or cannot be resolved. |
Evidence model
The reasoning package contains hypothesis versions, prediction records, test designs, admissible evidence, dependency groups, outcome classifications and revision history. A support statement must name the predicted feature and observed feature; a contradiction must name the tolerance exceeded. Missing evidence remains a node with planned acquisition, owner role and status.
Controlled workflow
- Define the decision and spatial, temporal and scale scope.
- Write multiple plausible hypotheses without inspecting holdout evidence.
- Derive observable predictions and explicit disconfirmation conditions.
- Map evidence lineage and discount dependent repetitions.
- Classify outcomes as support, contradiction, non-diagnostic or unknown.
- Select the next test by expected discrimination and consequence.
Every step emits a versioned artefact or an explicit failure. A later stage consumes only verified outputs from the preceding stage; conversational context is never an undocumented data channel.
Measures and acceptance criteria
Measure prediction specificity, test coverage, dependency-adjusted evidence coverage, contradiction handling and revision responsiveness. A useful matrix contains tests that could fail each hypothesis; a matrix of universally supportive statements has no discrimination. Record the fraction of decisive predictions evaluated and the number of unresolved assumptions that control the conclusion.
$D(t)=\sum_{i<j} w_{ij}\,\left|P(o_t\mid H_i)-P(o_t\mid H_j)\right|$
| Gate | Required evidence | Release consequence | |---|---|---| | Identity | Resolvable claim, source, configuration and result IDs | Block when any identity is ambiguous | | Grounding | Every factual claim reaches sufficient source spans | Remove or withhold unsupported claims | | Independence | Lineage and partition checks show no circular or target information | Invalidate affected support and scores | | Review | Required roles disposition the exact candidate version | Keep the candidate non-published | | Reproducibility | Manifest, tools and checks reconstruct the evidence package | Return the package for correction |
Claim–source and system contract
Each hypothesis record declares scope, assumptions, predicted observations, disconfirmation conditions, dependencies, authoring state and review version. Evidence scoring uses a frozen rubric and preserves raw classifications. The system may summarise the matrix but cannot replace unknown with neutral support or convert absence of evidence into contradiction without an observation model.
Security, privacy and access boundary
Keep holdout observations and outcome labels inaccessible while hypotheses and predictions are authored. Record who can reveal them and when. If sensitive evidence must be summarised, retain a protected full record and a visible redaction state; do not let hidden evidence silently change a public score.
Human review and escalation
Domain reviewers check plausibility and prediction meaning; an independent role checks that scoring was frozen before evidence reveal. Review focuses on the strongest contradiction and the most consequential unknown, not only the preferred hypothesis. Changes after reveal are new hypothesis versions and are not rescored as if prespecified.
Worked synthetic example
Two synthetic hypotheses explain a conductivity anomaly: H1 predicts a steep tabular body aligned with a mapped structure; H2 predicts a shallow weathering feature tied to topography. Existing conductivity supports both and is non-diagnostic. A synthetic depth slice supports H1, while a surface geochemistry pattern weakly supports H2 but shares the same gridding source as the anomaly. The matrix discounts that dependence and identifies one independent depth observation as the next decisive test.
Counterexample and failure analysis
A generated explanation lists five facts that fit H1 and concludes “high confidence” without stating H2, predicted failures or missing tests. The count is meaningless because four facts derive from one processed dataset. Structured reasoning reveals one lineage group, one non-diagnostic observation and several unknowns rather than five independent confirmations.
Frequent failure modes
- Generating one story before alternatives
- Writing predictions after seeing outcomes
- Counting dependent derivatives independently
- Treating missing evidence as support
- Hiding contradictions in narrative summaries
Practical exercise
- Write two hypotheses and three discriminating predictions for a synthetic anomaly.
- Classify six evidence items and justify every relation.
- Identify shared lineage and revise the apparent support count.
- Select a next test using discrimination, cost boundary and consequence.
Assessed artefact
Submit a reviewed competing-hypothesis matrix with frozen predictions and test plan with its source manifest, configuration versions, acceptance evidence, rejected cases, unresolved risks and a short explanation of why one plausible alternative control was not selected. The artefact is incomplete if its diagrams or prose cannot be reconciled with machine-checkable evidence.
Verification checkpoint
Trace one released synthetic claim backwards through review, confidence, verification, evidence graph, tool result or source span to an immutable input. Then trace one rejected claim to the earliest failed contract. Re-run the package from its manifest and compare content digests. The checkpoint passes only when a second reviewer can reconstruct the accepted and blocked paths, identify every assumption and reproduce the release decision without access to the original conversation. Record unresolved risk instead of completing a missing fact with generated text.
Sources
- Strong Inference, a primary account of designing decisive tests among alternative explanations.
- AI Risk Management Framework 1.0, a lifecycle framework for governing, mapping, measuring and managing context-dependent AI risk.
- PROV-O provenance ontology, a standard model for entities, activities, responsibility roles and derivation.
- Strictly Proper Scoring Rules, Prediction, and Estimation, a formal basis for evaluating probabilistic statements without rewarding strategic confidence.