GenoAtlas works from genomic data rather than proprietary or undocumented genetic-coordinate systems.
Genomic data
Modern reference individuals are primarily selected from publicly available whole-genome sequencing (WGS) datasets and genomic data contributed by donors. Ancient individuals are obtained from published archaeogenomic studies and public sequencing repositories.
Depending on the original study, genomic data may be available as BAM, CRAM, VCF, or sequencing reads. GenoAtlas processes these sources into a common representation so that ancient and modern individuals can be compared using the same marker set.
Reference populations are progressively expanded to provide broader geographic coverage while retaining individual-level genomic data rather than relying solely on population averages.
Data processing
A custom Java pipeline built with HTSJDK is used to read and process genomic files, extract genotypes, and harmonize genetic markers across datasets.
Because genomic datasets may have been generated using different sequencing technologies, coverage levels, and reference genome assemblies, GenoAtlas identifies positions that can be consistently compared within each collection.
Only the shared marker set retained for a particular collection is used for downstream population-genetic analyses. The number of markers used is displayed directly on each collection page.
Principal Component Analysis
Principal Component Analysis (PCA) is performed using PLINK 2.
The PCA space is defined using modern whole-genome reference individuals. Ancient individuals are subsequently projected into this reference space, allowing their genetic affinities to be visualized without allowing ancient samples to redefine the principal components themselves.
This provides a stable modern reference framework in which ancient individuals from different archaeological contexts can be compared.
Genetic distances
GenoAtlas provides individual-level comparisons between ancient samples and modern reference individuals using their positions in PCA space.
These distances are exploratory measures of genetic similarity. They do not represent direct ancestry percentages and should be interpreted together with the PCA distribution, archaeological context, and additional population-genetic analyses.
Y-DNA and mtDNA
Where sufficient information is available, GenoAtlas also displays Y-chromosome and mitochondrial DNA haplogroups.
These uniparental markers represent only the paternal or maternal lineage of an individual and therefore describe a very small part of total ancestry. They are presented alongside autosomal analyses rather than as substitutes for them.
Ancient haplogroup assignments derived computationally may be refined as additional sequence information or manual review becomes available.
Multiple analytical approaches
Each GenoAtlas collection represents a shared set of ancient and modern genomic samples rather than a single PCA analysis. Different analytical methods can therefore be applied to the same collection.
PCA · Y-DNA / mtDNA · qpAdm · ADMIXTURE
Availability depends on the collection and the current stage of analysis. Each method answers a different population-genetic question and should be interpreted within its own assumptions and limitations.
Transparency
GenoAtlas maintains a clear distinction between original published genomic data, GenoAtlas data processing, and results generated by the analytical tools used by the platform.
Original studies remain the authoritative source for archaeological context, sample dating, sequencing methodology, and primary scientific conclusions. GenoAtlas provides an independent environment for processing, comparing, and interactively exploring those publicly available genomic datasets.
Tools & Technologies
- PLINK 2
- PCA and population-genetic data processing
- HTSJDK
- Genomic BAM, CRAM, and VCF processing
- Custom GenoAtlas tools
- Genotype extraction, dataset harmonization, sample processing, and interactive analysis
- qpAdm / ADMIXTURE
- Additional population-genetic analyses where available
