RESULTS: Using two independent gene-prediction pipelines, Fgenesh++ and Seqping, 26,059 oil palm genes with transcriptome and RefSeq support were identified from the oil palm genome. These coding regions of the genome have a characteristic broad distribution of GC3 (fraction of cytosine and guanine in the third position of a codon) with over half the GC3-rich genes (GC3 ≥ 0.75286) being intronless. In comparison, only one-seventh of the oil palm genes identified are intronless. Using comparative genomics analysis, characterization of conserved domains and active sites, and expression analysis, 42 key genes involved in FA biosynthesis in oil palm were identified. For three of them, namely EgFABF, EgFABH and EgFAD3, segmental duplication events were detected. Our analysis also identified 210 candidate resistance genes in six classes, grouped by their protein domain structures.
CONCLUSIONS: We present an accurate and comprehensive annotation of the oil palm genome, focusing on analysis of important categories of genes (GC3-rich and intronless), as well as those associated with important functions, such as FA biosynthesis and disease resistance. The study demonstrated the advantages of having an integrated approach to gene prediction and developed a computational framework for combining multiple genome annotations. These results, available in the oil palm annotation database ( http://palmxplore.mpob.gov.my ), will provide important resources for studies on the genomes of oil palm and related crops.
REVIEWERS: This article was reviewed by Alexander Kel, Igor Rogozin, and Vladimir A. Kuznetsov.
RESULTS: In general, the genetic diversity decreased from Costa Rica towards the north (Honduras) and south-east (Colombia). Principle coordinate analysis (PCoA) showed a single cluster indicating low divergence among palms. The phylogenetic tree and STRUCTURE analysis revealed clusters based on country of origin, indicating considerable gene flow among populations within countries. Based on the values of the genetic diversity parameters, some genetically diverse populations could be identified. Further, a total of 34 individual palms that collectively captured maximum allelic diversity with reduced redundancy were also identified. High pairwise genetic differentiation (Fst > 0.250) among populations was evident, particularly between the Colombian populations and those from Honduras, Panama and Costa Rica. Crossing selected palms from highly differentiated populations could generate off-springs that retain more genetic diversity.
CONCLUSION: The results attained are useful for selecting palms and populations for core collection. The selected materials can also be included into crossing scheme to generate offsprings that capture greater genetic diversity for selection gain in the future.