Reaching the long tail of disease data means inferring. To score a disease that has never had a dedicated study, a ranking has to borrow: propagate signals from related diseases, read single-cell expression in the affected tissue, fall back to organ-level context. There is nothing wrong with that. It is the only way to say anything at all about disease biology the literature has barely touched.
The problem is not the inference. It is the presentation. A single priority rank number cannot tell you whether a call rests on human causal genetics or on an organ-level guess. In a system that reports a single priority number, two candidates can share that number for completely different reasons, one anchored in a genetic experiment, the other in a soft expression pattern. The figure looks identical; the ground beneath it is not. It is for this reason that Disease Atlas keeps the two apart.
Disease Atlas by Euretos attaches a confidence grade to every disease and phenotype, tied to the strength of the evidence that seeds it. The strongest grade is direct causal evidence: Mendelian randomisation on genetic effect sizes, with perturbation concordance and curated genetics alongside it. Below that sits supporting evidence, propagated or indirect. At the base is exploratory signal for the long tail, read from expression or resolved only to the organ.
The grade is not cosmetic. It propagates from the seed all the way through to the output, so a confident-looking score can never launder a weak signal into a strong claim. Coverage that carries its own confidence is a different product from coverage alone. One hands you a ranking. The other tells you which parts of the ranking you can build on today, and which parts are a hypothesis still to test.
We run this evidence grading across the whole atlas, and show the grade for every disease and phenotype. Roughly 34,000 diseases and phenotypes carry a target ranking. The tiers are nested: 81% carry direct or supporting evidence, and within that, 54% of all diseases sit in the higher-confidence tier. The remaining 19% are the exploratory long tail, flagged as such rather than dressed up as certainty.
Within the higher-confidence tier, one number matters most: 77% of those diseases rest on a direct causal-genetics seed. In plain terms, each of them has at least one target whose link to the disease comes from a Mendelian randomisation signal that is statistically strong and points in a consistent biological direction. Mendelian randomisation uses natural genetic variation to test whether a gene actually influences a disease, rather than simply appearing alongside it. So the high-confidence label is not a curation artefact. Roughly three-quarters of it traces back to causal human genetics.
By therapeutic area the pattern is informative rather than flat. Infectious disease (90%), respiratory (85%), ophthalmic (84%) and metabolic and endocrine (83%) are the most genetics-dense. The large neurology and oncology pools sit a little lower, near 72 to 74%, because their higher-confidence sets include more diseases resting on fine-mapped GWAS loci and curated evidence than on direct causal inference. Even there it is close to three-quarters. The chart with this piece shows both halves: how the confidence tiers stack up across every area, and how much of each area's high-confidence set is carried by causal genetics.
In common, well-mapped disease, a mis-graded call is recoverable. There are other signals, other groups working the same biology, other programmes to fall back on, so an answer graded too high or too low tends to get caught. In rare and long-tail disease the candidate pool is small, and a single wrong call can sink the effort. That is exactly where clear evidence grading earns its keep. It tells you when a lead is strong enough to advance and when to treat it as a question still to be answered.
And because the strongest grade is anchored in causal genetics, the diseases the Disease Atlas is most confident about tend to be the ones where the biology is most likely to hold. Targets with human genetic support are roughly twice as likely to reach approval. The grade that points to the strongest evidence is the same grade that raises the odds of success.
Disease Atlas takes a position on every disease and grades its own confidence. Any disease, from the most-studied to the far long tail, carries a defensible strength rating and a rationale that traces back to its source. The count is the easy headline. The grade is what lets a research team act on the count instead of guessing which part of it to believe.
Breadth without a confidence attached is a claim. Breadth that grades itself, disease by disease, down to whether the call rests on causal genetics, is a decision you can defend.