Research

Analysis of human genomes with genome graphs

The reference genome is a cardinal element in any genome study. It is a blueprint against which newly sequenced genomes get compared for read mapping, variant calling, and further analyses. The current widely used human reference genome (hg38) represents each chromosome by a linear, continuous string of nucleotide bases. hg38 in its linear form cannot capture the genetic information from different human populations and lead to reference allele bias.
The study focuses on constructing and using genome graphs as an alternate reference structure for genome research. Scalable computational pipelines were developed by integrating multiple bioinformatics tools to analyze human WGS with pan-genomic graphs. Novel methods of annotating the genome graph were designed to identify functionally significant regions. Detailed analyses were performed to study the structural complexities of the genome graphs.

GenomeIndia: Unveiling unique variants in Indian (sub)populations

Previous works have established that substantial differences are present in genomes from different human populations across the world. The Indian population is highly heterogeneous, with an estimate of over 4,500 genetically distinct sub-populations. Moreover, it is also divergent from the much-studied Caucasian genome. The GenomeIndia project aims to sequence the genomes of 10K healthy individuals from the country’s diverse populace and study them to catalog the common and low-frequency variants in Indians.
Genome graphs have the potential to detect novel variants in under-represented populations as they overcome reference allele bias. Our study focuses on constructing Indian population-specific genome graphs and using them as reference structures to analyze Indian genomes. The dynamic nature of the genome graphs enhances the detection of pathogenic variants in the Indian population. Functional annotation of Indian genome graphs helps understand the consequences of genetic variations within and between Indian (sub)populations.

Approachability of genomics and visualization of variants

Genomics is a rapidly evolving field that has profoundly impacted biomedical and public health research. As the applications of genomics burgeon into various domains, it calls for tools that empower clinicians and researchers to work with genomic data formats irrespective of their programming expertise.
We designed and developed SCI-VCF as a cross-platform application to summarise and compare the genetic variants from VCF files. Users can also create customizable interactive visualizations based on the contents of VCF files. The guided GUI setting of the tool empowers researchers and clinicians to perform exploratory genomic data analysis on VCF files irrespective of their programming expertise. The user-friendly and intuitive design of SCI-VCF increases the approachability of genomics to newcomers and introduces genomic data analysis expeditiously.

Polygenic Risk Scores for Common Complex Diseases

Complex diseases occur due to the combined effect of multiple genomic mutations at different loci, in conjunction with physical and environmental factors. Genome-wide association Studies (GWAS) statistically assess the strength of the relationship between genomic variants and complex diseases. Using the effect scores from GWASs, polygenic risk scores (PRS) quantify the risk for an individual for common complex diseases by calculating the cumulative effects of the genomic variants in an individual. PRS is most informative for early detection and is valuable when devising intervention and treatment strategies.
In our study, we built a computational pipeline to calculate PRS for type-2 diabetes for individuals in an Indian cohort. Different GWAS studies were experimented with as base datasets. Various ways to improve the predictive accuracy of the calculated PRS, like linkage disequilibrium treatments, clumping, and p-value thresholding, were explored. Machine learning techniques were employed to enhance the performance of the computational pipeline. This analysis lays the foundation for future studies to develop multi-ethnic and pathway-specific polygenic risk score predictions.

Understanding the Pathogenicity through the Lens of Genome Graphs

Understanding the genomic hotspots that drive the transition from asymptomatic carriage to invasive infection remains a challenge in infectious disease research. Traditional reliance on linear reference genomes limits genetic diversity representation and hinders in-depth research. In our study, we used genome graphs to investigate the genomic basis of pathogenicity in Streptococcus pneumoniae, a globally relevant bacterium with diverse lineages varying in pathogenicity. We analysed these microbial genome graphs using GViNC, a comprehensive framework for genome graph analyses. We conducted locus-specific analyses and applied graph distance metrics to identify key regions of variability differentiating infection from carriage isolates. Our approach offers a scalable framework for integrating genome graphs into routine microbial surveillance, aiding in the study of pathogen evolution in diverse environments.


More Information

Click on the respective tiles below to view a catalog of my publications, presentations, and training.