Skip to main content An official website of the United States government Here's how you know Here's how you know Official websites use .gov A .gov website belongs to an official government organization in the United States. Secure .gov websites use HTTPS A lock ( Lock Locked padlock icon ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites. Search Log in Dashboard Publications Account settings Log out Search… Search NCBI Primary site navigation Search Logged in as: Dashboard Publications Account settings Log in Search PMC Full-Text Archive Search in PMC Journal List User Guide PERMALINK Copy As a library, NLM provides access to scientific literature. Inclusion in an NLM database does not imply endorsement of, or agreement with, the contents by NLM or the National Institutes of Health. Learn more: PMC Disclaimer | PMC Copyright Notice Nat Commun . 2026 Mar 2;17:3342. doi: 10.1038/s41467-026-70175-y Search in PMC Search in PubMed View in NLM Catalog Add to search Fine-scale structure of a whole regional population through genetics and genealogies Gilles-Philippe Morin Gilles-Philippe Morin 1 Département des sciences fondamentales, Université du Québec à Chicoutimi, Saguenay, QC Canada 2 Centre Intersectoriel en Santé Durable (CISD), Université du Québec à Chicoutimi (UQAC), Saguenay, QC Canada 3 Unité mixte de recherche INRS-UQAC en santé durable, Saguenay, QC Canada Find articles by Gilles-Philippe Morin 1, 2, 3 , Claudia Moreau Claudia Moreau 1 Département des sciences fondamentales, Université du Québec à Chicoutimi, Saguenay, QC Canada 2 Centre Intersectoriel en Santé Durable (CISD), Université du Québec à Chicoutimi (UQAC), Saguenay, QC Canada Find articles by Claudia Moreau 1, 2 , Amadou Barry Amadou Barry 2 Centre Intersectoriel en Santé Durable (CISD), Université du Québec à Chicoutimi (UQAC), Saguenay, QC Canada 3 Unité mixte de recherche INRS-UQAC en santé durable, Saguenay, QC Canada 4 Centre Armand-Frappier Santé Biotechnologie, Institut national de la recherche scientifique (INRS), Laval, QC Canada Find articles by Amadou Barry 2, 3, 4 , Simon L Girard Simon L Girard 1 Département des sciences fondamentales, Université du Québec à Chicoutimi, Saguenay, QC Canada 2 Centre Intersectoriel en Santé Durable (CISD), Université du Québec à Chicoutimi (UQAC), Saguenay, QC Canada 3 Unité mixte de recherche INRS-UQAC en santé durable, Saguenay, QC Canada 5 Projet BALSAC, Université du Québec à Chicoutimi (UQAC), Saguenay, QC Canada 6 Centre de recherche CERVO, Université Laval, Québec, QC Canada Find articles by Simon L Girard 1, 2, 3, 5, 6, ✉ Author information Article notes Copyright and License information 1 Département des sciences fondamentales, Université du Québec à Chicoutimi, Saguenay, QC Canada 2 Centre Intersectoriel en Santé Durable (CISD), Université du Québec à Chicoutimi (UQAC), Saguenay, QC Canada 3 Unité mixte de recherche INRS-UQAC en santé durable, Saguenay, QC Canada 4 Centre Armand-Frappier Santé Biotechnologie, Institut national de la recherche scientifique (INRS), Laval, QC Canada 5 Projet BALSAC, Université du Québec à Chicoutimi (UQAC), Saguenay, QC Canada 6 Centre de recherche CERVO, Université Laval, Québec, QC Canada ✉ Corresponding author. Received 2025 Oct 16; Accepted 2026 Feb 20; Collection date 2026. © The Author(s) 2026 Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/ . PMC Copyright notice PMCID: PMC13066195 PMID: 41771909 Abstract Population stratification can confound genetic association studies and often persists despite adjustment using principal components of common variants. Demographic history and rare variants can also contribute to this confounding. But to what extent does demographic history impact fine structure and can be detected in human populations? To address this question, we analysed the Saguenay–Lac-Saint-Jean region of Quebec, a recent founder population long assumed to be homogeneous. Integrating genotype data with genealogical records, we show a strong concordance between realised (genetic) and expected (genealogical) kinship. Using a time-efficient algorithm capable of computing billions of pairwise kinship coefficients for all individuals married in the region between 1931 and 1960, we reveal fine structure at the municipal level, including an east–west genetic gradient shaped by differential founders’ contribution, migration patterns and socioeconomic factors. These findings challenge the assumption of regional homogeneity and suggest that similar recent structure likely exists in other populations worldwide. These signals may be obscured by coarse, ancient structure under standard stratification corrections used in genome-wide association studies and polygenic risk score analyses which can lead to biases and false associations when realised allele frequencies correlate with phenotypic variation. Subject terms: Population genetics, Population genetics, Consanguinity, Genome informatics Population stratification can bias genetic studies, yet fine recent structure is often overlooked. Here, the authors integrate genotypes and genealogies from Saguenay-Lac‑Saint‑Jean, revealing municipal‑level structure shaped by founders, migration and socioeconomic history. Introduction Population stratification occurs when genetic structure (non-random allele distribution) and environmental differences between subpopulations confound genetic association studies. This can make environmental effects appear genetic, leading to misleading results in genome-wide association studies (GWAS) and polygenic scores 1 , 2 . Family-based GWAS, which controls for a shared familial environment, reveals that standard population GWAS often overestimates genetic effects 3 . This provides evidence that population stratification may persist despite conventional population structure controls, challenging assumptions in large-scale GWAS. When these GWAS are used to construct polygenic scores, even small biases can accumulate into large errors. Critically, it has been shown that the impact of population stratification is also shaped by demographic history which is often overlooked even after adjusting for 100 principal components of common variants 4 . This study suggests that using principal components derived from rare variants, which capture most of the genetic diversity, or from identity-by-descent (IBD) segments may correct the stratification for some types of effects. IBD analysis has proven effective for uncovering fine-scale population structure in large cohorts. For example, a study using the UK Biobank 5 revealed subtle genetic substructure among British individuals, while research on New York City residents identified 16 ancestry groups, including seven founder populations 6 , through the All of Us Research Program and the Mount Sinai BioMe biobank. Fine-scale structure has also been documented in the Quebec founder population 7 – 9 . These findings raise an important question: to what extent can fine-scale structure be detected in populations? In this study, we address this question by examining a small, recent founder population, the Saguenay–Lac-Saint-Jean (SLSJ) region, leveraging extensive genotype and genealogical data to characterise its genetic architecture. The SLSJ region of Quebec, Canada (Fig. 1 ), represents a remarkable model for population genetics research due to its well-documented founder effect, recent settlement and unique demographic history. From the seventeenth century, European settlers of mostly French origin have migrated along the Saint Lawrence River, then founded the Charlevoix region, before settling in Saguenay and Lac-Saint-Jean from 1838 10 . This serial founder effect was accentuated by a 25-fold population surge in SLSJ within a century, mainly due to high birth rates 11 , 12 . The region’s demographic characteristics are documented through the BALSAC population register 13 , a comprehensive database documenting civil records, primarily Catholic marriages, of French-Canadian individuals from the 1660 s (New France) to the 1960s (Quebec), encompassing over three centuries of demographic history. This demographic history is known to have led to several variants being more frequent in the region 14 and to the increase of some rare diseases 15 – 18 . Some researchers have characterised the genetic background of SLSJ as relatively homogeneous 16 , 19 , suggesting that, unlike the broader provincial context, the region might lack meaningful structure. Yet, demographic studies indicate that population structure does exist within SLSJ 20 , 21 , challenging the assumption of homogeneity. Fig. 1. Map of Saguenay–Lac-Saint-Jean municipalities. Open in a new tab Circle size reflects the number of probands married between 1931–1960 in each BALSAC municipality (total 80,348). The map was generated by the authors using public domain cartographic data from Natural Earth and river geometries from OpenStreetMap (© OpenStreetMap contributors), available under the Open Database License (ODbL). Previous research in the French-Canadian population has shown that both genotype and genealogical data correlate substantially 7 – 9 , 22 , 23 . This shows that genealogy may be used as a proxy to study the genetic structure of populations. Yet, despite the wealth of genealogical and demographic data available for the SLSJ region, significant challenges remain in adequately visualising the region’s population structure at a large scale. This limitation primarily stems from algorithms that repeatedly compute kinship for redundant pairs of individuals within a genealogy 24 , making analysis of complete populations computationally intractable. Although efforts have optimised these computations 25 , 26 , implementing these methods at the scale needed for whole-population analysis, particularly for visualisation purposes, has remained challenging, restricting samples to hundreds or a few thousand individuals at a time. The absence of appropriate analytical methods has thus limited comprehensive investigation of fine-scale structure in populations. This study aims to comprehensively investigate the fine-scale population structure within the SLSJ region by integrating genotype and genealogical datasets. We also retrace the population’s origins through the expected genetic contribution of the region’s founders. By doing so, this article reveals a fine-scale population structure shaped by geographical boundaries, differential founders’ genetic contribution and socioeconomic factors in the recent SLSJ founder population. Results Comparison between genetic and genealogical structure The CARTaGENE cohort comprises individuals with both genotype and genealogical data, of which 7970 probands have sufficient genealogical completeness and are not siblings (see ‘Methods’). Uniform Manifold Approximation and Projection (UMAP) with two dimensions based on realised and expected kinship (see ‘Methods’) display strikingly similar population structure patterns (Fig. 2 ); a linear regression between realised and expected kinship further supports this consistency with a Pearson correlation coefficient of 0.78 ( p < 0.001). Hierarchical Density-Based Spatial Clustering of Applications with Noise (HDBSCAN) on a UMAP with ten dimensions identified a subset of 938 individuals most likely originating from the SLSJ region (Supplementary Fig. 1 ). Within this subset, UMAP projections derived from both realised and expected kinship (Fig. 3 ) reveal partially overlapping gradients of individuals from the diverse subdivisions of SLSJ (Supplementary Fig. 2 ) with a significant correlation between realised and expected kinship of 0.83 ( p < 0.001). Fig. 2. Two-dimensional UMAP of CARTaGENE individuals based on pairwise realised and expected kinship. Open in a new tab Uniform Manifold Approximation and Projection (UMAP) of 7970 individuals from the CARTaGENE cohort, computed from their ( a ) realised kinship and ( b) expected kinship transformed as a precomputed distance ( 1 − ϕ ). The colours represent the region of parents’ marriage. Fig. 3. Two-dimensional UMAP of CARTaGENE individuals from SLSJ based on pairwise realised and expected kinship. Open in a new tab Uniform Manifold Approximation and Projection (UMAP) projection of 938 individuals from the CARTaGENE cohort who are most likely to originate from Saguenay–Lac-Saint-Jean (SLSJ), computed from their ( a ) realised kinship and ( b ) expected kinship transformed as a precomputed distance ( 1 − ϕ ). The colours represent the SLSJ watercourse subdivision of parents’ marriage. Fine-scale population structure of a whole generation The great correlation between the population structure observed through genetic and genealogical data 7 – 9 , 22 , 23 allows us to use the genealogy as a proxy for examining the fine-scale structure of the whole SLSJ population of the last generation available in BALSAC. The two-dimensional UMAP projection of non-siblings probands married between 1931 and 1960 in the SLSJ region reveals clear concentrations of individuals whose parents were married in the same subdivision (Fig. 4 ). The observed structure follows an east–west gradient, extending from the southeast (East of Ha! Ha! subdivision) to the northwest (Ashuapmushuan–Mistassini subdivision) (see Fig. 1 ). This gradient exhibits a circular pattern centred around individuals originating from Charlevoix and concludes with a distinct grouping of individuals whose ancestry lies outside both Charlevoix and SLSJ. More remote, rural municipalities exhibit higher levels of concentration in genealogical relationships, while urban centres display greater dispersion (Supplementary Fig. 3 ). Fig. 4. Two-dimensional UMAP of pairwise expected kinship for the last-generation SLSJ population. Open in a new tab Uniform Manifold Approximation and Projection (UMAP) of 26,445 non-siblings who married between 1931 and 1960 in Saguenay–Lac-Saint-Jean (SLSJ), computed from their expected kinship ( ϕ ) transformed as a precomputed distance (1 − ϕ ). The colours represent the watercourse subdivision of parents’ marriage. Expected genetic contribution of founders To refine our understanding of the migratory patterns underlying the observed structure, we calculated the expected genetic contribution of the SLSJ founders (ancestors married in SLSJ whose parents married elsewhere, see ‘Methods’) to the SLSJ probands of the last generation. Mean genetic contribution from the SLSJ founders originating from Charlevoix is the highest in the eastern part of SLSJ and progressively declines along an east–west axis (Fig. 5a ) 21 . Individuals from the northwesternmost municipalities of Lac-Saint-Jean exhibit more diverse ancestral origins, as well as individuals originating from urban areas, especially south of Saguenay River that exhibit high diversity. A detailed analysis (Fig. 5b, c ) reveals that this east–west genetic gradient is primarily driven by founders from La Malbaie and to some extent Les Éboulements, which both contributed most significantly to the Saguenay area (see also Supplementary Fig. 4 ). In contrast, Baie-Saint-Paul contributed more to the Lac-Saint-Jean region (Fig. 5b, c , see also Supplementary Fig. 4 ). Notably, Les Éboulements contributed disproportionately to the east of Ha! Ha! which also presents the highest contribution from Charlevoix founders overall (Supplementary Fig. 5 ). Individuals from the area between Ashuapmushuan and Mistassini largely trace their ancestry to regions outside Charlevoix (Fig. 5b, c ), with this diversification becoming prominent around 1885 (Supplementary Fig. 5 ). Fig. 5. Proportion of the mean expected genetic contribution to probands explained by SLSJ founders depending on geography and on various Charlevoix municipalities. Open in a new tab Mean expected genetic contribution of Saguenay–Lac-Saint-Jean (SLSJ) founders from Charlevoix to 80,348 probands as a function of the longitude of the probands’ subdivisions ( a ). The red line shows the unweighted linear regression with a Pearson correlation coefficient of 0.517 ( p < 0.001). Mean expected genetic contribution of SLSJ founders from Charlevoix per municipality to probands’ in each subdivision ( b , c ). The map was generated by the authors using public domain cartographic data from Natural Earth and river geometries from OpenStreetMap (© OpenStreetMap contributors), available under the Open Database License (ODbL). Discussion In this study, we found fine-scale population structure in SLSJ, challenging the notion of homogeneity in this founder population 16 , 19 . This fine-scale structure was assessed and visualised on the entire SLSJ population, thanks to deep genealogical data and a time-efficient hybrid algorithm for computing kinship coefficients on such a massive scale. It is therefore reasonable to expect that a fine structure due to recent demographic history exists in all human populations. As already known, the fine-scale population structure may have an impact on rare variants 27 , 28 and association studies 29 , 30 . Quebec is known for regional enrichments of rare variants, most notably in SLSJ 14 , but evidence points to similar patterns in other less studied regions 28 , 31 – 33 . Our study suggests that such enrichments may be more localised than at the regional scale, even within regions with a strong founder effect. Within the population of a single region, the carrier rate may vary from one location to another. If a variant is more frequent in one part of the region and less in another, considering only the overall carrier rate may mitigate the assessment of enrichment and lead to an under-appreciation of the variant’s frequency in the region. This factor is intensified if the carrier tests are performed unevenly, such as in urban centres more than in rural municipalities. Beyond influencing the distribution of rare variants, fine-scale population structure may pose challenges for GWAS, especially when it results in allele frequencies that correlates with phenotypic variations 34 , 35 . This could introduce confounding factors that standard correction methods may not fully address. Previous findings suggesting polygenic adaptation on height across European populations were later shown to be largely attributable to uncontrolled stratification, prompting a reassessment of polygenic score comparisons between populations 1 , 2 . Subsequent work has demonstrated that geographic structure remains detectable in a large homogeneous cohort even after adjustment for study centre and genetic principal components 36 . Using simulations, we demonstrated that the SLSJ fine-scale population structure, shaped by heterogeneous founder contributions and migration routes, can induce detectable stratification in a GWAS that PCA of common SNPs fails to completely account for (Supplementary Fig. 6 and Supplementary Methods). Detecting such fine-scale structure within a relatively small and homogeneous population underscores the critical need to control for population stratification in GWAS and polygenic analyses beyond the first 20 or 100 principal components generally used 4 . In this study we reaffirm the correlation between genetics and genealogy 7 – 9 , 22 , 23 . Therefore, using genealogy as a proxy for genetic structure is valuable given its strong correlation with realised kinship, which has been shown to capture fine-scale population structure in other populations 5 , 6 . Indeed, it enabled us to study the fine structure of the entire SLSJ population, which would not be possible with genetic data alone, which are unavailable for all individuals and are often concentrated in urban areas, highlighting the need for inclusion of smaller and remote communities in future genetic studies 37 – 40 . This work also showcases some methodological advances, such as the power of parallelised kinship computation via GeneaKit: billions of coefficients inferred efficiently across tens of thousands of individuals. Our hybrid pipeline, melding Karigl’s 24 and Kirkpatrick’s 26 algorithms, proved to be time‑efficient and is readily extensible to other deep genealogy datasets worldwide. By applying the hybrid kinship algorithm to the whole SLSJ population, it allowed us to uncover a fine population structure which may be partly explained by the migratory patterns that shaped the SLSJ population. Previous work described a Charlevoix gradient, in which eastern municipalities show a higher genetic contribution from Charlevoix founders than those farther west 21 . However, we show here for the first time at a finer scale that this gradient is less homogeneous than suggested. Indeed, it centres on La Malbaie, the main contributor to Saguenay, while Baie‑Saint‑Paul’s influence predominates in Lac-Saint-Jean. In the eastern part of the region, Les Éboulements shows a strong but previously undocumented genetic contribution to the East of Ha! Ha! subdivision, following an early wave of migration from La Malbaie. Moving westward, individuals show less than half of their genetic input from Charlevoix, instead tracing largely to external Quebec regions other than Charlevoix. In the northwest of Lac-Saint-Jean, the opening of the railroad in 1888 10 brought in diverse founders from outside Charlevoix 21 . By 1912, many Acadians settled south of the Saguenay River attracted by jobs in the paper and aluminium industries 41 , diluting Charlevoix’s relative genetic contribution in those urban centres. Indeed, the SLSJ fine structure was shaped not only by geography and migration but also by socioeconomic factors. While our study provides valuable insights into Quebec’s genetic structure, it is important to acknowledge certain limitations. The CARTaGENE dataset, although rich in genotype data, primarily comprises individuals recruited in urban areas, which might lead to an underrepresentation of more isolated communities. Similarly, the BALSAC genealogies, while extensive, exhibit regional variations in completeness 8 . This uneven data quality could potentially introduce biases when detecting population structure. Furthermore, our use of parents’ marriage location appears to be a better proxy for regional affiliation than the recruitment locations of CARTaGENE. It may still not fully account for individual or family mobility over generations or situations where parents originate from different regions. Despite these limitations, our rigorous quality control, which involved filtering out individuals with more than half null kinship coefficients, allowed us to observe the Quebec-wide genetic structure previously described. This resolution was sufficient to distinguish the peripheries of the SLSJ region. We also found that our UMAP projections, despite the inherent stochastic nature of the algorithm that can make direct comparisons challenging 42 , yielded highly similar results with matrices of expected and realised kinship. These findings underscore the robustness of our analysis in capturing the underlying genetic landscape of SLSJ and the province of Quebec. In conclusion, integrating large‑scale genealogical and genotype data reveals fine‑scale structure in SLSJ that broadly aligns with geographic gradients. Such genetic heterogeneity suggests uneven distributions of founder variants and, by extension, locally variable prevalence of associated genetic diseases. These findings also have implications for how we understand human genetic diversity more broadly. The use of self-reported race and ethnicity as proxies for human genetic diversity has long been widespread, and recent applications of UMAP have been employed to support such categorisations 43 . In reality, population structure and genetic diversity are observable not only at the continental scale, but also within regional populations, such as those in the province of Quebec. Our findings further demonstrate that population structure can be detected at even finer geographic resolutions, down to the level of municipalities, where we observe gradients of genetic contribution from different founders in addition to sharply defined clusters for more remote localities. Such regional and local gradients can impact genetic association studies, and should not be overlooked. Methods This study was approved by the ethics board of the Université du Québec à Chicoutimi (UQAC) and complies with all relevant ethical regulations. Genotype data and cleaning Genotype data was drawn from the CARTaGENE cohort 44 , which comprises 43,032 participants aged 40–69 (in 2005) recruited in the main urban areas of the Quebec Province, notably in the city of Saguenay. Of those participants, 29,337 individuals were genotyped using a variety of platforms, including Omni 2.5, GSAv1 with a multi-disease panel, GSAv1, GSAv2 with a multi-disease panel, GSAv3 with a multi-disease panel, GSAv2 with a multi-disease panel and add-on (see https://cartagene.qc.ca/files/documents/other/Info_GeneticData3juillet2023.pdf for details). Each dataset was processed independently using PLINK v1.9 45 , retaining individuals with at least 95% genotyping across all SNPs. At the SNP level, we retained autosomal variants with a call rate of at least 95% across individuals and that conformed to Hardy–Weinberg equilibrium ( p > 10 −6 , calculated within each dataset). All chips were merged, and the final dataset comprised 148,200 SNPs across 28,358 individuals. Genealogical data Genealogical data was obtained from the BALSAC population register 13 . Two different genealogical datasets have been extracted for this study: First, 9405 Quebec residents from the CARTaGENE cohort have been matched to BALSAC records. Second, to analyse a whole generation, we also selected 80,348 probands (i.e. non-parent individuals) married in SLSJ from 1931 to 1960. For both probands’ sets, a multigenerational genealogy extending up to 19 generations was reconstructed. When available, information on the municipality, region and year of marriage was included; to ensure confidentiality, marriage years are rounded up to the nearest 5. Individual anonymity is rigorously maintained through the use of unique, non-identifying codes. Traditional geographic subdivisions split SLSJ into Lower Saguenay, Upper Saguenay and Lac-Saint-Jean 20 . However, it has been demonstrated that geographic features and natural barriers, such as watersheds, can influence patterns of population structure 7 . We carefully carved eight subdivisions along primary watercourses to consolidate smaller communities and municipalities were grouped based on these watercourse boundaries (Supplementary Table 1 and Fig. 1 ). On average, we observed reduced migration in the last generation within our defined geographical subdivisions than between them (Supplementary Fig. 7 ). IBD sharing (realised kinship) Pairwise IBD segments were inferred on phased genotypes using Refined IBD (version 17Jan20) 46 and Beagle (version 18May20) 47 . A matrix of realised kinship estimation was computed by summing shared segment lengths of at least 2 centiMorgans (cM) using a Python 3 ported version of relatedness_v1.py from Sharon R. Browning (see https://faculty.washington.edu/sguy/ibd_relatedness.html ). The diagonal was set to 1. Genealogical kinship coefficients (expected kinship) Expected kinship ( ϕ ) 48 was used to infer and visualise population structure through genealogical data. It corresponds to the probability that, at one locus, one randomly picked allele from individual i and one randomly picked allele from individual j are IBD or come from the same ancestor. Expected kinship coefficients were computed between all pairs of probands from whole SLSJ and CARTaGENE genealogical datasets. Of the 80,348 SLSJ probands, 26,445 are non-siblings ( ϕ > 0.2) and have a genealogy deemed sufficiently complete, i.e. they are related ( ϕ > 0) with at least half of the probands. Of the 9405 individuals in CARTaGENE data, 7970 are neither parents nor siblings, and have a genealogy deemed sufficiently complete using the same criterion. Hybrid algorithm for time-efficient computation of billions of kinship coefficients To efficiently compute kinship coefficients across all pairs of individuals, we implemented a hybrid of two existing algorithms. Karigl’s algorithm 24 estimates kinship coefficients on a per-proband-pair basis, enabling parallel processing of the kinship matrix, as in the GENLIB package for R 49 . However, it reprocesses each ancestor every time they appear, either multiple times across genealogies (kinship) or repeatedly within the same genealogy (inbreeding), which makes it computationally costly. Kirkpatrick’s algorithm 26 addresses this inefficiency by partitioning the genealogy into successive sub-genealogies (e.g. one generation at a time) and computing kinship coefficients in a top-down manner, with the probands of one sub-genealogy treated as the founders of the next, thereby minimising redundant ancestor processing. Yet, Kirkpatrick’s approach is not designed for parallel execution. Our method combines the strengths of both: we divide the genealogy into generational sub-genealogies, as in Kirkpatrick’s approach, but compute each sub-genealogy’s kinship matrix in parallel, achieving substantial time gains. For instance, computing kinship for all 6,455,801,104 pairs of 1931–1960 BALSAC SLSJ probands required only 2 min with our algorithm. It was not possible to compare directly with Kirkpatrick’s implementation due to the fact that this software is no longer available online, nor with the GENLIB implementation because its execution would require more resources than can be allocated. Following the pre-publication of this work, it was brought to our attention that an indirect method by Colleau 50 can also be applied to compute kinship coefficients at a large scale 51 . However, this method incorporates all generations in the genealogy, which was not the objective of our analysis. Our implemented hybrid algorithm is provided as part of GeneaKit 0.1.0, which is freely available on GitHub 52 ( https://github.com/Genopop/geneakit ) under the MIT license, using the phi() function for kinship coefficients. A pseudocode of the algorithm is available in the supplementary material (Supplementary Note 1 ). Expected genetic contributions We also computed the expected genetic contribution of the region’s founders as a way to measure the region’s migratory origins and to better understand the fine structure. Region’s founders are defined as individuals who married in the SLSJ whereas their parents married elsewhere. This computation was done using GeneaKit 0.1.0’s gc() function, from all regions’ founders to all 80,348 genealogical probands who married between 1931 and 1960 in the SLSJ region. The expected genetic contribution of an ancestor corresponds to the probability that an allele is passed down to a given descendant. A founder’s region of origin is defined as the region of marriage of their parents. Clustering and visualisation CARTaGENE individuals originate from all regions of Quebec, it is thus necessary to identify those who are most likely to originate from SLSJ. UMAP (version 0.5.9) 42 , 53 was performed on both the realised and expected kinship matrices, using precomputed distances of ( 1 − ϕ ). UMAP was run with ten dimensions, 15 neighbours, a minimum distance of 0.0 and a fixed random seed (random state = 0), to maximise clustering and reproducibility. UMAP was initialised from a classical multidimensional scaling (MDS), also known as principal coordinate analysis (PCoA), of the precomputed distances, implemented as pcoa from scikit-bio 0.6.3 54 , using eigendecomposition and a seed of 0. HDBSCAN from scikit-learn 1.8.0 55 was then fitted to the UMAP outputs to identify individuals most likely originating from the SLSJ region, using a min_cluster_size of 25 and a cluster_selection_epsilon of 0.3. Those parameters were chosen as they allowed separation of regional populations in a previous study using the same cohort 44 . From the clustering results for both the realised and expected kinship matrices, the cluster containing the highest number of individuals of known SLSJ origin (i.e. whose four grandparents married in the region) was identified. 585/1141 individuals in the realised kinship cluster and 583/1036 individuals in the expected kinship cluster are of known SLSJ origin. The final set of 938 individuals was then determined as the intersection of these two best clusters (one from the realised kinship matrix clustering and one from the expected kinship matrix clustering). Of those 938 individuals, 566 are of known SLSJ origin. For visualisation, UMAP (initialised with classical MDS/PCoA) was applied to the kinship matrices using their precomputed distances ( 1 − ϕ ), two dimensions and a minimum distance of 0.5. Reporting summary Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article. Supplementary information Supplementary Information (5.2MB, pdf) Reporting Summary (88.8KB, pdf) Transparent Peer Review file (3.5MB, pdf) Acknowledgements This work was made possible by the Digital Research Alliance of Canada which provided access to storage and computing resources. We are extremely grateful to all participants in this research. Funding for S.L.G. was provided by the Canada Research Chair in Genetics and Genealogy CRC-2022-00444, of which he is the chairholder ( http://www.chairs.gc.ca ). Funding for G.-P.M. was provided by the Unité mixte de recherche INRS-UQAC en santé durable (grant number UIU0002). Author contributions All authors approved the final version of the manuscript. G.-P.M. and C.M. played an important role in interpreting the results. G.-P.M., C.M. and S.L.G. conceived and designed the study and drafted the manuscript. G.-P.M. performed the analyses. A.B. revised the manuscript. Peer review Peer review information Nature Communications thanks Hanbin Lee and the other, anonymous, reviewer(s) for their contribution to the peer review of this work. A peer review file is available. Data availability Due to the informed consent given by study participants, Quebec genotype data is available under restricted access via CARTaGENE’s independent Sample and Data Access Committee (SDAC) ( https://cartagene.qc.ca/en/researchers/access-request.html ). Further information on access to CARTaGENE data is available in the article by McClelland et al. 44 . Access to the BALSAC genealogical data is controlled in accordance with the Policy on Access to BALSAC Data for Research Purposes to protect participant confidentiality and adhere to ethical guidelines. The procedure for requesting access to BALSAC data is described at https://balsac.uqac.ca/acces-donnees/ . Access requests must be submitted to the Researchers Service of the BALSAC Project and must include a completed BALSAC Database Access Request Form along with all documents necessary for evaluation, such as a certificate of ethics approval. Applications are reviewed to ensure compliance with the scientific method, feasibility and BALSAC access policies. Applicants are informed of the decision in writing. Questions regarding data access can be addressed to [email protected]. If access is granted, data may be used only within the scope of the approved research project and must be destroyed once the authorised access period expires. The BALSAC data are not publicly available due to ethical restrictions on the publication of genealogical information dating back less than 100 years. Code availability The GeneaKit library is available under the open-source MIT license on GitHub ( https://github.com/Genopop/geneakit ) and is archived on Zenodo (10.5281/zenodo.18701153) 52 . The source code used for data processing, analysis and the production of figures in this study has been deposited in the following GitHub repository: https://github.com/Genopop/slsj_structure . It is also available on Zenodo: 10.5281/zenodo.18701179 56 . Competing interests The authors declare no competing interests. Footnotes Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations. Supplementary information The online version contains supplementary material available at 10.1038/s41467-026-70175-y. References 1. Sohail, M. et al. Polygenic adaptation on height is overestimated due to uncorrected stratification in genome-wide association studies. eLife 8 , e39702 (2019). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 2. Berg, J. J. et al. Reduced signal for polygenic adaptation of height in UK Biobank. eLife 8 , e39725 (2019). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 3. Tan, T. et al. Family-GWAS reveals effects of environment and mating on genetic associations. MedRxiv Preprint Server for Health Sciences 10.1101/2024.10.01.24314703 (2025). 4. Zaidi, A. A. & Mathieson, I. Demographic history mediates the effect of stratification on polygenic scores. eLife 9 , e61548 (2020). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 5. Nait Saada, J. et al. Identity-by-descent detection across 487,409 British samples reveals fine scale population structure and ultra-rare variant associations. Nat. Commun. 11 , 6130 (2020). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 6. Isshiki, M. et al. Genetic disease risks of under-represented founder populations in New York City. PLoS Genet. 21 , e1011755 (2025). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 7. Anderson-Trocmé, L. et al. On the genes, genealogies, and geographies of Quebec. Science 380 , 849–855 (2023). [ DOI ] [ PubMed ] [ Google Scholar ] 8. Gagnon, L., Moreau, C., Laprise, C., Vézina, H. & Girard, S. L. Deciphering the genetic structure of the Quebec founder population using genealogies. Eur. J. Hum. Genet . 32 , 91–97 (2024). [ DOI ] [ PMC free article ] [ PubMed ] 9. Gauvin, H. et al. Genome-wide patterns of identity-by-descent sharing in the French Canadian founder population. Eur. J. Hum. Genet. 22 , 814–821 (2014). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 10. Gagnon, S. Sagamiens—Sagamiennes. Saguenayensia 30 , 3–6 (1988). [ Google Scholar ] 11. Pouyez, C., Lavoie, Y. & Bouchard, G. Les Saguenayens: introduction à l’histoire des populations du Saguenay, XVIe-XXe siècles (Presses de l’Université du Québec, 1983). 12. Bouchard, G.& De Braekeleer, M. Histoire d’un génôme: Population et génétique dans l’est du Québec (Presses de l’Université du Québec, 1991). 13. BALSAC. https://balsac.uqac.ca/ . 14. Michel, É. et al. Rare diseases load through the study of a regional population. PLOS Genet. 21 , e1011876 (2025). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 15. Cruz Marino, T. et al. Portrait of autosomal recessive diseases in the FRENCH-CANADIAN founder population of SAGUENAY-LAC-SAINT-JEAN. Am. J. Med. Genet. A 191 , 1145–1163 (2023). [ DOI ] [ PubMed ] [ Google Scholar ] 16. Bchetnia, M. et al. Genetic burden linked to founder effects in Saguenay–Lac-Saint-Jean illustrates the importance of genetic screening test availability. J. Med. Genet. 58 , 653–665 (2021). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 17. Laberge, A. et al. Population history and its impact on medical genetics in Quebec. Clin. Genet. 68 , 287–301 (2005). [ DOI ] [ PubMed ] [ Google Scholar ] 18. Scriver, C. R. Human genetics: lessons from Quebec populations. Annu. Rev. Genom. Hum. Genet 2 , 69–101 (2001). [ DOI ] [ PubMed ] [ Google Scholar ] 19. Cruz Marino, T. et al. First glance at the molecular etiology of hearing loss in French-Canadian families from Saguenay-Lac-Saint-Jean’s founder population. Hum. Genet. 141 , 607–622 (2022). [ DOI ] [ PubMed ] [ Google Scholar ] 20. Lavoie, E.-M., Tremblay, M., Houde, L. & Vézina, H. Demogenetic study of three populations within a region with strong founder effects. Public Health Genom. 8 , 152–160 (2005). [ DOI ] [ PubMed ] [ Google Scholar ] 21. De Braekeleer, M. L'approche des maladies héréditaires par la démographie génétique: le cas du Saguenay-Lac-Saint-Jean au Québec . PhD thesis, Univ. Bordeaux 2 (1995). 22. Roy-Gagnon, M.-H. et al. Genomic and genealogical investigation of the French Canadian founder population structure. Hum. Genet. 129 , 521–531 (2011). [ DOI ] [ PubMed ] [ Google Scholar ] 23. Burkett, K. M. et al. Correspondence between genomic- and genealogical/coalescent-based inference of homozygosity by descent in large French-Canadian genealogies. Front. Genet. 12 , 808829 (2022). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 24. Karigl, G. A recursive algorithm for the calculation of identity coefficients. Ann. Hum. Genet. 45 , 299–305 (1981). [ DOI ] [ PubMed ] [ Google Scholar ] 25. Abney, M. A graphical algorithm for fast computation of identity coefficients and generalized kinship coefficients. Bioinformatics 25 , 1561–1563 (2009). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 26. Kirkpatrick, B., Ge, S. & Wang, L. Efficient computation of the kinship coefficients. Bioinformatics 35 , 1002–1008 (2019). [ DOI ] [ PubMed ] [ Google Scholar ] 27. Zhang, B. C., Biddanda, A., Gunnarsson, ÁF., Cooper, F. & Palamara, P. F. Biobank-scale inference of ancestral recombination graphs enables genealogical analysis of complex traits. Nat. Genet 55 , 768–776 (2023). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 28. Gagnon, L., Moreau, C., Laprise, C. & Girard, S. L. Fine-scale genetic structure and rare variant frequencies. PLoS ONE 19 , e0313133 (2024). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 29. Henn, B. M., Gravel, S., Moreno-Estrada, A., Acevedo-Acevedo, S. & Bustamante, C. D. Fine-scale population structure and the era of next-generation sequencing. Hum. Mol. Genet. 19 , R221–R226 (2010). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 30. Hu, S. et al. Fine-scale population structure and widespread conservation of genetic effect sizes between human groups across traits. Nat. Genet. 57 , 379–389 (2025). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 31. Gagnon, M. et al. Rare variants and founder effect in the Beauce region of Quebec. Commun. Biol. 8 , 1184 (2025). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 32. Vézina, H. et al. Molecular and genealogical characterization of the R1443X BRCA1 mutation in high-risk French-Canadian breast/ovarian cancer families. Hum. Genet. 117 , 119–132 (2005). [ DOI ] [ PubMed ] [ Google Scholar ] 33. De Braekeleer, M., Hechtman, P., Andermann, E. & Kaplan, F. The French Canadian Tay-Sachs disease deletion mutation: identification of probable founders. Hum. Genet. 89 , 83–87 (1992). [ DOI ] [ PubMed ] [ Google Scholar ] 34. Edge, M. D., Gorroochurn, P. & Rosenberg, N. A. Windfalls and pitfalls. Evol. Med. Public Health 2013 , 254–272 (2013). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 35. Rosenberg, N. A. & Nordborg, M. A general population-genetic model for the production by population structure of spurious genotype–phenotype associations in discrete, admixed or spatially distributed populations. Genetics 173 , 1665–1678 (2006). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 36. Haworth, S. et al. Apparent latent structure within the UK Biobank sample has implications for epidemiological analysis. Nat. Commun. 10 , 333 (2019). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 37. Brown, J. T., McGonagle, E., Seifert, R., Speedie, M. & Jacobson, P. A. Addressing disparities in pharmacogenomics through rural and underserved workforce education. Front. Genet. 13 , 1082985 (2023). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 38. Cohen, A. S. A. et al. Genomic answers for kids: toward more equitable access to genomic testing for rare diseases in rural populations. Am. J. Hum. Genet 111 , 825–832 (2024). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 39. Best, S., Vidic, N., An, K., Collins, F. & White, S. M. A systematic review of geographical inequities for accessing clinical genomic and genetic services for non-cancer related rare disease. Eur. J. Hum. Genet. 30 , 645–652 (2022). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 40. Fatumo, S. et al. A roadmap to increase diversity in genomic studies. Nat. Med 28 , 243–250 (2022). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 41. Gouvernement du Québec. Place des Acadiens—Saguenay (Ville). Comm. Topon. https://toponymie.gouv.qc.ca/ct/ToposWeb/Fiche.aspx?no_seq=446160 . 42. McInnes, L., Healy, J. & Melville, J. UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction. 10.48550/ARXIV.1802.03426 (2018). 43. The All of Us Research Program Genomics Investigators et al. Genomic data in the All of Us Research Program. Nature 627 , 340–346 (2024). [ DOI ] [ PMC free article ] [ PubMed ] 44. McClelland, P. et al. A multi-ancestry genetic reference for the Quebec population. Nat. Commun. 17 , 1319 (2026). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 45. Purcell, S. et al. PLINK: a tool set for whole-genome association and population-based linkage analyses. Am. J. Hum. Genet. 81 , 559–575 (2007). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 46. Browning, B. L. & Browning, S. R. Improving the accuracy and efficiency of identity-by-descent detection in population data. Genetics 194 , 459–471 (2013). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 47. Browning, S. R. & Browning, B. L. Rapid and accurate haplotype phasing and missing-data inference for whole-genome association studies by use of localized haplotype clustering. Am. J. Hum. Genet 81 , 1084–1097 (2007). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 48. Wright, S. The interpretation of population structure by F-statistics with special regard to systems of mating. Evolution 19 , 395 (1965). [ Google Scholar ] 49. Gauvin, H. et al. GENLIB: an R package for the analysis of genealogical data. BMC Bioinform. 16 , 160 (2015). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 50. Colleau, J.-J. An indirect approach to the extensive calculation of relationship coefficients. Genet Sel. Evol. 34 , 409 (2002). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 51. Lee, H., Craddock, R. F., Gorjanc, G. & Becher, H. randPedPCA: rapid approximation of principal components from large pedigrees. Genet Sel. Evol. 57 , 46 (2025). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 52. Morin, G. P. et al. Genopop/geneakit: geneakit 0.1.0. 10.5281/ZENODO.18701153 (2026). 53. McInnes, L., Healy, J. & Melville, J. UMAP: Uniform Manifold Approximation and Projection. J. Open Source Softw. 3 , 861 (2018). 54. Aton, M., McDonald, D., Cañardo Alastuey, J. et al. Scikit-bio: a fundamental Python library for biological omic data analysis. Nat. Methods 23 , 274–276 (2026). [ DOI ] [ PubMed ] 55. Pedregosa, F. et al. Scikit-learn: machine learning in Python. J. Mach. Learn Res 12 , 2825–2830 (2011). [ Google Scholar ] 56. Morin, G. P. et al. Fine-scale structure of a whole regional population through genetics and genealogies . 10.5281/zenodo.18701179 (2026). [ DOI ] [ PMC free article ] [ PubMed ] Associated Data This section collects any data citations, data availability statements, or supplementary materials included in this article. Supplementary Materials Supplementary Information (5.2MB, pdf) Reporting Summary (88.8KB, pdf) Transparent Peer Review file (3.5MB, pdf) Data Availability Statement Due to the informed consent given by study participants, Quebec genotype data is available under restricted access via CARTaGENE’s independent Sample and Data Access Committee (SDAC) ( https://cartagene.qc.ca/en/researchers/access-request.html ). Further information on access to CARTaGENE data is available in the article by McClelland et al. 44 . Access to the BALSAC genealogical data is controlled in accordance with the Policy on Access to BALSAC Data for Research Purposes to protect participant confidentiality and adhere to ethical guidelines. The procedure for requesting access to BALSAC data is described at https://balsac.uqac.ca/acces-donnees/ . Access requests must be submitted to the Researchers Service of the BALSAC Project and must include a completed BALSAC Database Access Request Form along with all documents necessary for evaluation, such as a certificate of ethics approval. Applications are reviewed to ensure compliance with the scientific method, feasibility and BALSAC access policies. Applicants are informed of the decision in writing. Questions regarding data access can be addressed to [email protected]. If access is granted, data may be used only within the scope of the approved research project and must be destroyed once the authorised access period expires. The BALSAC data are not publicly available due to ethical restrictions on the publication of genealogical information dating back less than 100 years. The GeneaKit library is available under the open-source MIT license on GitHub ( https://github.com/Genopop/geneakit ) and is archived on Zenodo (10.5281/zenodo.18701153) 52 . The source code used for data processing, analysis and the production of figures in this study has been deposited in the following GitHub repository: https://github.com/Genopop/slsj_structure . It is also available on Zenodo: 10.5281/zenodo.18701179 56 . Articles from Nature Communications are provided here courtesy of Nature Publishing Group ACTIONS View on publisher site PDF (4.1 MB) Cite Collections Permalink PERMALINK Copy RESOURCES Similar articles Cited by other articles Links to NCBI Databases Cite Copy Download .nbib .nbib Format: AMA APA MLA NLM Add to Collections Create a new collection Add to an existing collection Name your collection * Choose a collection Unable to load your collection due to an error Please try again Add Cancel Follow NCBI NCBI on X (formerly known as Twitter) NCBI on Facebook NCBI on LinkedIn NCBI on GitHub NCBI RSS feed Connect with NLM NLM on X (formerly known as Twitter) NLM on Facebook NLM on YouTube National Library of Medicine 8600 Rockville Pike Bethesda, MD 20894 Web Policies FOIA HHS Vulnerability Disclosure Help Accessibility Careers NLM NIH HHS USA.gov Back to Top