Jakun, a Proto-Malay subtribe from Peninsular Malaysia, is believed to have inhabited the Malay Archipelago during the period of agricultural expansion approximately 4 thousand years ago (kya). However, their genetic structure and population history remain inconclusive. In this study, we report the genome structure of a Jakun female, based on whole-genome sequencing, which yielded an average coverage of 35.97-fold. We identified approximately 3.6 million single-nucleotide variations (SNVs) and 517,784 small insertions/deletions (indels). Of these, 39,916 SNVs were novel (referencing dbSNP151), and 10,167 were nonsynonymous (nsSNVs), spanning 5674 genes. Principal Component Analysis (PCA) revealed that the Jakun genome sequence closely clustered with the genomes of the Cambodians (CAM) and the Metropolitan Malays from Singapore (SG_MAS). The ADMIXTURE analysis further revealed potential admixture from the EA and North Borneo populations, as corroborated by the results from the F3, F4, and TreeMix analyses. Mitochondrial DNA analysis revealed that the Jakun genome carried the N21a haplogroup (estimated to have occurred ~19 kya), which is commonly found among Malays from Malaysia and Indonesia. From the whole-genome sequence data, we identified 825 damaging and deleterious nonsynonymous single-nucleotide polymorphisms (nsSNVs) affecting 720 genes. Some of these variants are associated with age-related macular degeneration, atrial fibrillation, and HDL cholesterol level. Additionally, we located a total of 3310 variants on 32 core adsorption, distribution, metabolism, and elimination (ADME) genes. Of these, 193 variants are listed in PharmGKB, and 21 are nsSNVs. In summary, the genetic structure identified in the Jakun individual could enhance the mapping of genetic variants for disease-based population studies and further our understanding of the human migration history in Southeast Asia.
* Title and MeSH Headings from MEDLINE®/PubMed®, a database of the U.S. National Library of Medicine.