Transparent and reproducible population-genetics workflow

Methodology

GenoAtlas processes public genomic data with established population-genetics software and custom harmonization tools.

GenoAtlas works from genomic data rather than proprietary or undocumented genetic-coordinate systems.

Genomic data

Modern reference individuals are primarily selected from publicly available whole-genome sequencing (WGS) datasets and genomic data contributed by donors. Ancient individuals are obtained from published archaeogenomic studies and public sequencing repositories.

Depending on the original study, genomic data may be available as BAM, CRAM, VCF, or sequencing reads. GenoAtlas processes these sources into a common representation so that ancient and modern individuals can be compared using the same marker set.

Reference populations are progressively expanded to provide broader geographic coverage while retaining individual-level genomic data rather than relying solely on population averages.

Data processing

A custom Java pipeline built with HTSJDK is used to read and process genomic files, extract genotypes, and harmonize genetic markers across datasets.

Because genomic datasets may have been generated using different sequencing technologies, coverage levels, and reference genome assemblies, GenoAtlas identifies positions that can be consistently compared within each collection.

Only the shared marker set retained for a particular collection is used for downstream population-genetic analyses. The number of markers used is displayed directly on each collection page.

Principal Component Analysis

Principal Component Analysis (PCA) is performed using PLINK 2.

The PCA space is defined using modern whole-genome reference individuals. Ancient individuals are subsequently projected into this reference space, allowing their genetic affinities to be visualized without allowing ancient samples to redefine the principal components themselves.

This provides a stable modern reference framework in which ancient individuals from different archaeological contexts can be compared.

Genetic distances

GenoAtlas provides individual-level comparisons between ancient samples and modern reference individuals using their positions in PCA space.

These distances are exploratory measures of genetic similarity. They do not represent direct ancestry percentages and should be interpreted together with the PCA distribution, archaeological context, and additional population-genetic analyses.

Y-DNA and mtDNA

Where sufficient information is available, GenoAtlas also displays Y-chromosome and mitochondrial DNA haplogroups.

These uniparental markers represent only the paternal or maternal lineage of an individual and therefore describe a very small part of total ancestry. They are presented alongside autosomal analyses rather than as substitutes for them.

Ancient haplogroup assignments derived computationally may be refined as additional sequence information or manual review becomes available.

Multiple analytical approaches

Each GenoAtlas collection represents a shared set of ancient and modern genomic samples rather than a single PCA analysis. Different analytical methods can therefore be applied to the same collection.

PCA · Y-DNA / mtDNA · qpAdm · ADMIXTURE

Availability depends on the collection and the current stage of analysis. Each method answers a different population-genetic question and should be interpreted within its own assumptions and limitations.

Transparency

GenoAtlas maintains a clear distinction between original published genomic data, GenoAtlas data processing, and results generated by the analytical tools used by the platform.

Original studies remain the authoritative source for archaeological context, sample dating, sequencing methodology, and primary scientific conclusions. GenoAtlas provides an independent environment for processing, comparing, and interactively exploring those publicly available genomic datasets.

Tools & Technologies

PLINK 2
PCA and population-genetic data processing
HTSJDK
Genomic BAM, CRAM, and VCF processing
Custom GenoAtlas tools
Genotype extraction, dataset harmonization, sample processing, and interactive analysis
qpAdm / ADMIXTURE
Additional population-genetic analyses where available