Catalog failure domains
Assign validators to organizations, regions, infrastructure providers, and other correlated-risk groups.
THE IDEA, MADE VISIBLE
Group validators by an infrastructure domain, then remove entire domains instead of independent machines.
Six validator processes may look like six independent units of resilience. Their deployment relationships determine whether that count is meaningful.
Losing 1 of 3 equally sized hosting domains leaves 4 validators. The chosen availability predicate survives this scenario.
Correlated failures act on domains, not independent validator counts.
Choose one domain and fail it. All six validators disappear together, even though the machine count never changed.
What this experiment represents. Constructed hosting topology and availability check only. These are not actual Metriq infrastructure records, a full fault-domain policy search or a Byzantine-safety certificate.
Adversary policies, domain-cover margins, and exact portfolio certification
Certify quorum safety under correlated failures rather than validator counts alone. Organization, region and infrastructure incidents are modeled explicitly through domain-cover policies and exact witness searches.

The paper certifies quorum safety against correlated failures—such as one organization, region, or infrastructure provider failing—rather than pretending validators fail independently.
Ten validators do not provide ten independent sources of resilience if eight share one cloud provider or legal operator. Ordinary validator counts can therefore exaggerate safety.
The proposed method declares failure domains and adversary policies explicitly, then searches for quorum pairs and domain covers that expose the smallest correlated failure capable of breaking intersection.
The paper’s technical details matter, but the basic route can be understood in three moves.
Assign validators to organizations, regions, infrastructure providers, and other correlated-risk groups.
Define which combinations of domains may fail together and what counts as an admissible incident.
Compute quorum pairs and minimum domain covers that certify or refute the desired safety margin.
Correlated failure is modeled directly instead of approximated through validator count.
The output includes exact witnesses for weak points and domain-cover margins for certified cases.
The supplied topology is an illustrative case study, not evidence about production Metriq infrastructure.
Fault-domain certification methods and exact analysis of an illustrative topology. The case study is not production Metriq infrastructure evidence.
This report develops a conditional method and evaluates it within the stated evidence. It does not establish production performance, operational safety, or calibrated real-world predictive skill beyond that evidence.
A certificate is only as truthful as the declared failure-domain catalog.
The case study does not establish production-network resilience.
Unmodeled common causes can invalidate a seemingly strong margin.
Counting validators can badly overstate resilience when many share one failure domain. Explicit domain-cover policies expose those correlated risks.
These are the terms needed to understand the claim. The full paper uses them more precisely.
A group of components that may fail together because they share an operator, region, provider, or dependency.
A failure event affecting several validators for one shared reason.
The formal set of failure combinations a certification promises to tolerate.
A concrete quorum or fault configuration demonstrating that a claimed margin fails.
The strongest review is not a general reaction. It tests the steps most capable of changing the conclusion.
This page is a reading guide, not a substitute for the manuscript. The public record links the explanation to the paper, source package, review materials, and persistent identifier.