Define admissible evidence
Specify timing, spatial support, source version, and endpoint rules for matching predictions to observations.
THE IDEA, MADE VISIBLE
Uncertain observations should widen a comparison—not let each model choose its own favorable reality.
Model A predicts 0.35 and model B predicts 0.70 for one constructed normalized outcome. These are demonstration numbers, not fitted reef predictions.
Using the same y, error A minus error B lies in [-0.0525, 0.1575]. The ranking is not determined over the admitted interval.
Δ(y) = (0.35−y)² − (0.70−y)²
Widen the radius until the paired-error interval crosses zero. Neither model then wins under every admissible observation.
What this experiment represents. Exact formulas evaluated in floating point for a single constructed observation interval. No biological records, fitted models or empirical predictive advantage are being compared.
Sharp comparison bounds, source compatibility, and an executable chronology benchmark
Compare bleaching models using the same admitted observation alignment. The method supplies sharp finite comparison bounds, compatibility rules and reproducible synthetic benchmarks without claiming a fitted biological validation.

The paper says two bleaching models should be judged against the same admissible observation alignment, not each model’s individually most favorable version of the data.
Bleaching observations may cover different dates, locations, survey methods, and source versions. If each model is allowed to choose its own alignment, the comparison can be biased before any score is computed.
The proposed framework defines which alignments are compatible, forces both models onto the same admitted evidence, and reports sharp performance bounds when more than one alignment remains possible.
The paper’s technical details matter, but the basic route can be understood in three moves.
Specify timing, spatial support, source version, and endpoint rules for matching predictions to observations.
Evaluate competing models on the same compatible assignment rather than separate model-specific assignments.
Compute best- and worst-case score differences across the shared compatibility set and test the implementation on constructed benchmarks.
Observation alignment becomes part of the evaluation protocol rather than a hidden preprocessing choice.
Shared alignment prevents a model from winning by selecting easier evidence.
Sharp finite bounds make unresolved alignment uncertainty visible.
Validation methods, exact constructed comparisons, and computational benchmarks. Independent bleaching observations have not been acquired or fitted in this edition.
This report develops a conditional method and evaluates it within the stated evidence. It does not establish production performance, operational safety, or calibrated real-world predictive skill beyond that evidence.
No real-world model ranking is established without independent bleaching observations.
Dependence grouping and finite-support assumptions affect the validity of the bounds.
A fair evaluation protocol does not guarantee that either model is biologically correct.
A model should not win by choosing a more favorable alignment than its competitor. Shared admissible observations make the comparison itself auditable.
These are the terms needed to understand the claim. The full paper uses them more precisely.
The rule that matches a prediction to a particular observed outcome in time and space.
The time period, spatial area, or population represented by a measurement.
All prediction–observation matchings allowed by the stated evidence rules.
A controlled test case used to check whether an evaluation method behaves as expected.
The strongest review is not a general reaction. It tests the steps most capable of changing the conclusion.
This page is a reading guide, not a substitute for the manuscript. The public record links the explanation to the paper, source package, review materials, and persistent identifier.