Transparency
Disclaimer, data and sources
Where every number on this site comes from, how confident it is, what version it belongs to, and what DNADojo explicitly does not claim.
Updated 2026-09-13 · 6 min read · Genealogy & education only
What this site is for
DNADojo is a genealogy and education product. It helps you interpret shared-DNA numbers in terms of family relationships, and it helps you look inside a raw genotype file you already own. That is the whole scope.
Not medical information
Not legal or forensic evidence
Output from this site must not be used as evidence, or as the basis for a decision, in legal proceedings, parentage or paternity determination, immigration or citizenship applications, criminal or forensic investigation, insurance, employment, credit or housing decisions, or to identify a person who has not consented to being identified. Those uses require accredited testing under a documented chain of custody. See section 4 of the Terms of Service.
The relationship dataset
Every cM figure on this site comes from one versioned dataset file. The current version is v1.0.0-simulated, built 2026-09-14, containing 33 relationship records.
How the numbers are derived
- A published genetic map. Chromosome lengths come from Bhérer et al. 2017 refined sex-averaged genetic map (build 37). The import script reads the published files directly and records the SHA-256 of each one, so every figure here traces back to exact bytes rather than to numbers somebody retyped. The sex-averaged autosomal map totals 3,346.3 cM.
- Simulated meiosis. For each relationship we build the pedigree and drop crossovers along each chromosome as a Poisson process at one per 100 cM (Haldane, no interference). Founders carry unique labels, so any interval two people share is identical by descent by construction — there is no matching heuristic to get wrong.
- Reported the way a testing company reports it. Shared cM is the total length of half-identical regions, counting a fully identical region once. Segments below 7 cM are dropped, which mirrors real detection thresholds. This is why full siblings come out near 2,600 cM rather than half of the genome.
- One disclosed rescaling. A genetic map spans only the region its markers cover, which is slightly shorter than the scale testing companies report against. Every simulated value is multiplied by ×1.0414 so that parent/child reads 3,485 cM — the number you would see on your own results. That is a single constant applied to the whole distribution, not a per-record adjustment: the shape of every range, and the relationship between relationships, is exactly what the simulation produced.
- The displayed range is the widest and narrowest values seen across the simulated pairs that shared anything, which is how published tables are built too. Typical is the median of those pairs. How often a pair shares nothing at all is reported separately, as share any DNA.
- Identical twins are the one record not simulated. Twins share the whole genome by definition rather than through recombination, so that figure is anchored to reported values.
How it compares to published averages
Nothing here is derived from these figures — they are an external check. Close relationships land within a few percent:
| Relationship | DNADojo typical | Commonly published average | Difference | |
|---|---|---|---|---|
| Parent / child | 3,485 cM | Always | 3,485 cM | +0.0% |
| Full sibling | 2,594 cM | Always | 2,613 cM | -0.7% |
| Half sibling | 1,718 cM | Always | 1,759 cM | -2.3% |
| Grandparent / grandchild | 1,733 cM | Always | 1,754 cM | -1.2% |
| Aunt/uncle — niece/nephew | 1,705 cM | Always | 1,741 cM | -2.1% |
| Great-aunt/uncle — great-niece/nephew | 835 cM | Always | 851 cM | -1.9% |
| First cousin | 829 cM | Always | 866 cM | -4.3% |
| Second cousin | 194 cM | Always | 229 cM | -15.3% |
| Third cousin | 46 cM | 89.5% | 73 cM | -37.0% |
| Fourth cousin | 19 cM | 45.9% | 35 cM | -45.7% |
Every range on this site is conditional on the two people sharing detectable DNA. That is deliberate. A published shared-cM table can only contain pairs who matched — if two fourth cousins share nothing, neither of them is ever submitted to it — and a visitor holding a cM figure has, by definition, already matched. So the two are comparable, and the number answers the question actually being asked.
The unconditional fact is reported separately, as share any DNA on each relationship page, because for distant cousins it matters more than the range does. Two fourth cousins have roughly a 46% chance of sharing anything at all; sixth cousins almost never do. Not matching is not evidence that you are unrelated.
Our distant-cousin figures still run below published tables, and the gap widens with distance. We have not tuned that away, because the honest explanation is a bias we cannot remove: a published table contains relationships someone recognised and submitted, which skews toward larger matches, and real platforms apply matching thresholds and post-processing stricter than our flat 7 cM floor. Read a published distant-cousin average as the high end of what you might see, not the middle.
Sources
| ID | Source | Status | Summary |
|---|---|---|---|
genetic-map-v1 | Bhérer et al. 2017 refined sex-averaged genetic map (build 37) | VERIFIED | Bhérer C, Campbell CL, Auton A. Refined genetic maps reveal sexual dimorphism in human meiotic recombination at multiple scales. Nature Communications 8:14994 (2017). doi:10.1038/ncomms14994. Built from 3.3 million crossovers observed across 104,246 meioses. The sex-averaged autosomal map totals 3,346.3 cM. DNADojo reads the published files directly and records the SHA-256 of each one in src/data/genetic-map.json, so every centimorgan figure on this site can be traced to exact bytes. Shared-cM distributions are then produced by simulating 50,000 pedigrees per relationship along this map (Poisson crossovers at 1 per 100 cM, Haldane, no interference; 7 cM minimum segment; fully-identical regions counted once), and rescaling the result by a single disclosed constant so parent/child reads 3,485 cM, the value consumer testing companies report. |
provisional-model-v0 | DNADojo provisional Mendelian model (placeholder) | SUPERSEDED | Retired on 2026-09-14, replaced by genetic-map-v1. Expected shared percentages are the textbook coefficient of relationship. Centimorgan values are that percentage applied to a 6,800 cM working autosomal total, with a heuristic spread that widens with the number of meioses. NOT an empirical study. Must be replaced before launch. |
raw-data-informational-use | Consumer raw genotype downloads are provided for informational use | TO_VERIFY | Used on platform pages to state that raw genotype downloads are not a diagnostic product. Link the current vendor help-centre article before publishing. |
genetic-data-sensitive-category | Genetic data is treated as a sensitive data category by US regulators | TO_VERIFY | Supports the privacy page. Cite the specific regulator publication and date before publishing. |
anchored-reported-value | Anchored to values reported across consumer testing platforms | ANCHORED | Used for the identical-twin record only. Identical twins share the whole genome by definition rather than through recombination, so there is no distribution to simulate. The figure shown is roughly twice a parent/child result, matching what testing companies report. |
Changelog
| Version | Date | Change |
|---|---|---|
v1.0.0-simulated | 2026-09-14 | Replaced the model-derived placeholder set. Centimorgan values now come from 50,000 simulated pedigrees per relationship along the Bhérer et al. 2017 sex-averaged genetic map, rescaled by a single disclosed constant (x1.0414) so parent/child reads 3,485 cM. Displayed ranges are the extremes observed across the simulated pairs. Identical twin remains anchored. Still awaiting subject-matter review, so every record is marked 'simulated' and the provisional badge stays up. |
v0.1.0-provisional | 2026-09-13 | Initial provisional dataset generated by scripts/generate-provisional-dataset.mjs. 33 relationship records. Not reviewed. Not for production. |
Known limits
- Endogamy is not modelled. In populations with long histories of intermarriage, shared cM runs well above these ranges for the same stated relationship.
- Multiple relationships are not combined. If two people are related through more than one line, the true sharing is higher than any single label predicts.
- Platform differences are not modelled. Testing companies use different chips, genetic maps and matching thresholds.
- The X chromosome is not included. X-DNA has its own inheritance pattern and is out of scope for the current tools.
- Compatibility scores are not probabilities. The ranking shows relative fit across a shortlist; it is not a posterior probability and does not incorporate a prior.
Found something wrong? Corrections to the dataset are the most valuable contribution anyone can make to this site, and every change is recorded in the changelog above with its version.