E3 ยท Publication Volume 25
Interval Quality Rules
gaps, overlaps, negative length and intervals beyond end of hole
Learning objectives
- Explain the decision and evidence boundary for gaps, overlaps, negative length and intervals beyond end of hole.
- Design and implement the relevant drillhole data or algorithm contract without hidden conventions.
- Separate hard release gates from diagnostics, interpretation and authorised review.
- Produce an interval-QC engine and coverage report with exact affected spans from synthetic evidence.
The lesson is complete only when the learner can defend the data model, algorithm, tests and release decision. An attractive trajectory or clean interval table without source evidence and executable invariants remains unverified.
This is a general, institution-neutral tutorial with no relationship to any company or individual. All borehole identifiers, coordinates, depths, directions, intervals, values and review events in the lesson are synthetic and must not be used for an operational decision.
Decision context
Quality control decides whether an interval collection satisfies the rules of its observation type and intended use. A partition table may require continuous non-overlapping coverage; a sample table may legitimately contain gaps; independent interpretation tables may overlap by design. The rule set therefore declares collection type, expected depth extent, tolerance policy and whether each condition is invalid, warning or informational.
Write the intended use, consequence of error, required evidence and release authority before selecting a transformation. The same source can be suitable for exploratory display and unsuitable for a released derivative. Fitness is evaluated against a versioned contract and use, not attached permanently to a file.
Core concept
Validate each row before analysing relationships: finite boundaries, correct unit, positive length, resolved borehole and support within the allowed depth range. Then stable-sort by canonical from, to and identity, and sweep through endpoints to find gaps, overlaps, containment and duplicate spans. Findings identify the exact span and all contributing records. A tolerance classifies near-equal boundaries but does not erase the original discrepancy.
Keep received observations, accepted evidence views and derived results as distinct objects. This separation allows corrected evidence or a changed method to generate a new result without rewriting history. Every derived coordinate or interval therefore answers both a scientific question and a provenance question.
Algorithm and data model
Separate row findings from collection findings. Negative length and non-finite endpoints belong to one row. A gap belongs to a collection and an expected coverage extent. An overlap belongs to two or more records and may have several active contributors. Beyond-end-of-hole findings require a versioned end-of-hole observation and its depth datum. If that observation is absent or uncertain, report the dependency rather than inventing a limit.
Define the transformation as a pure, testable operation wherever practical. Parsing, semantic validation, evidence selection, numeric calculation and release evaluation are separate stages. Each stage emits structured output and does not depend on interface state, filename order or an undocumented default.
Constraints and invariants
| Invariant | Executable or review test | | --- | --- | | Collection expectations are declared before gap or overlap tests. | Reject or quarantine any record that violates this condition and record the exact affected identity. | | Coverage uses interval unions rather than summed row lengths. | Evaluate this condition before producing a derived trajectory or interval result. | | Every finding records the exact affected span. | Preserve received evidence and create a new version for every correction. | | End-of-hole checks reference a versioned observation. | Include the rule identifier, observed value and resolution state in audit output. |
An invariant must survive import, conversion, processing, export and rerun. A failed hard invariant produces no apparently valid substitute. Diagnostic checks remain visible with their threshold, scope and evidence, and require a reviewed rule before they can trigger correction.
Quantitative reasoning
For expected extent L, report union coverage C=|\cup_i I_i|/L, total gap length, union overlap length and multiplicity by span. Do not sum interval lengths to estimate coverage because overlaps are then counted repeatedly. During a sweep, events at equal depth follow the declared half-open ordering so interval ends are processed before starts. Tolerance-sensitive results include both raw discrepancy and classified state.
Every reported metric includes units, numerator and denominator where applicable, exclusions, comparison policy and evaluation version. Aggregate values are stratified when pooling could hide a local failure. A quantitative diagnostic supports a decision but cannot overrule missing identity, invalid geometry, unresolved conflict or broken lineage.
Evidence and uncertainty
Keep observation uncertainty, interpolation uncertainty, numeric approximation and metadata uncertainty separate. A smooth trajectory can be numerically precise while still poorly constrained between widely spaced stations. An exact interval overlay can still be unfit when a source depth datum is unknown. The assessed result states which uncertainty belongs to the phenomenon, the measurement, the algorithm and the interpretation.
Build an evidence packet containing immutable received records, semantic declarations, validation findings, algorithm inputs and outputs, test results, reviewer decisions and fingerprints. Contradictory evidence remains available. When a required dependency cannot be resolved, return an explicit unknown, conflict or blocked status rather than choosing the most convenient value.
Interfaces and storage
Interfaces transmit identities, units, coordinate and depth references, conventions, value states, versions and lineage beside numeric values. A trajectory exchange includes collar and datum context, accepted station identities, algorithm identity, numerical policy and output coordinates. An interval exchange includes support type, boundary convention and source links. Structured errors identify the record, field, observed value, expected condition and rule.
Store authoritative received evidence separately from reproducible derivatives and disposable views. Indexes, caches and visualisations may improve access but cannot become the only copy of angle conventions, accepted-station decisions or interval lineage. Export round trips verify that identifiers, precision, ordering and missing states survive encoding changes.
Governance and review
Assign responsibilities to roles rather than named organisations or people: evidence custodian, rule author, implementation maintainer, independent validator and release reviewer. A role may propose a correction but cannot erase source evidence. Rule and algorithm changes are reviewed, versioned and evaluated against fixed regression fixtures before they affect a release.
Exceptions are explicit decisions with scope, rationale, evidence, approving role, affected versions and review trigger. They never rewrite a failed rule and never propagate automatically. The host website has no ownership or scientific-authority role in this workflow; it only delivers the tutorial.
Integration checkpoint
Read the figure as a reasoning map from preserved evidence through explicit conventions, deterministic calculation, validation and release. Each arrow represents a declared relationship or transformation. Integrate an interval-QC engine and coverage report with exact affected spans into the evolving synthetic drillhole package, rerun all earlier fixtures and record any changed assumption.
Synthetic worked example
A synthetic partition intended to cover [0,100) contains [0,40), [39.9998,70), [72,90), [90,101). With a depth-comparison tolerance of 0.001 m, the first boundary discrepancy is classified as near-equal but retained. The exact [70,72) gap is a hard partition failure. The final interval exceeds the verified 100 m end of hole by 1 m and is blocked. Coverage is calculated from the union, not the 101.0002 m sum of row lengths.
- Preserve the received records and state the intended decision without correction.
- Resolve identities, units, conventions and evidence eligibility; mark every unresolved item.
- Run the versioned algorithm and tests while retaining intermediate diagnostics.
- Issue accept, reject or quarantine and show how an independent reviewer can reproduce it.
Practice task
Implement the chapter artefact against a synthetic fixture containing one normal case, one boundary case, one invalid case and one unresolved-evidence case. Preserve the received fixture. Produce canonical input, validation findings, derivative output, processing manifest and a short release decision.
Acceptance criteria:
- Every input identity, unit and convention required by the rule is explicit.
- The implementation is deterministic under stable ordering and the declared numerical policy.
- No correction overwrites received evidence or turns unknown into a guessed value.
- All hard failures block the affected derivative and remain machine-readable.
- A second implementation or reviewer can reproduce the result from the package alone.
Submit an interval-QC engine and coverage report with exact affected spans, the golden and adversarial fixtures, exact findings and a limitations note. A screenshot is not sufficient evidence because it does not identify the input version, algorithm or rule configuration.
Common failure modes
- Calling every gap an error regardless of collection type.
- Snapping boundaries before preserving raw differences.
- Summing overlap lengths multiple times.
- Assuming a maximum interval endpoint is end of hole.
These failures share a pattern: an implicit convenience is substituted for evidence. Diagnose the earliest boundary where the assumption entered, restore the source statement, make the convention or rule explicit, rerun every dependent derivative and supersede rather than overwrite the affected release.
Review questions
- Why must interval expectations be type-specific?
- How does a sweep-line algorithm find exact overlap spans?
- Why is summed row length not coverage?
- What evidence is required for an beyond-end-of-hole finding?
For every answer, identify the governing invariant, the evidence needed to evaluate it, the numerical or semantic policy involved and the correct behaviour when the condition fails.
Sources and further reading
- OGC GeoSciML 4.1 Borehole requirements, including borehole intervals, one-dimensional support and interval ordering.
- ISO 19157-1:2023 geographic data quality, a framework for describing and evaluating data quality.
- IEEE 754-2019 floating-point arithmetic, specifying floating-point formats, operations, rounding and exception behaviour.
- W3C PROV-O, a model for entities, activities, responsibility roles and derivation.