Variant ID in NyuWa Global 10K T2T Variant Database (NyuWaT2T)

The format of variant in NyuWaT2T is Chromosome-Position-ReferenceAllele-AlternativeAllele, for example, 13-47266103-C-T. The position, reference allele and alternative allele of variants are left-aligned and normalized (https://genome.sph.umich.edu/wiki/Variant_Normalization). The position coordinate of NyuWaT2T is based on human assembly T2T-CHM13.

Search result

After submitting the search key, there will be a feedback table which contains variants matching the querying. From the table, some overview information of the variants can be glanced, including Variant ID, dbSNP ID, information resulting from RefSeq gene annotation, and statisctical information from 10,457 high depth genome sequencing (Allele Number, Allele Frequency).


Special Projects

You can directly click on these projects from the homepage.

T2T-specific regions

Provide browsing and querying of variant information in T2T-specific genomic regions.

Imputation Webserver

Offer online imputation services based on a panel based of whole T2T reference genome. You can upload your own set of variants and perform imputation for various types of variants.

Regulatory Function

You can query functional annotation information here, which includes annotation results of xQTL and annotation results based on GWAS.

xQTL:


GWAS:

Population Differentiation

You can search for the allele frequency distribution of variant sites across various populations, here you can filter the result by variant type.

Linkage Disequilibrium

You can search Linkage disequilibrium (LD) between SNP/InDels and complex variants in the NyuWa T2T resource here.

Genotype Correction

You can find the allele frequency (AF) differences of variants between the T2T reference genome and the hg38 reference genome here. This platform also allows for quick searches of gene regions related to diseases or medical conditions.

Variant annotation page

Basic information

From the search result, every variant can be linked to a detail variant annotation page which is modularized. First it’s the basic information of variant including Allele Count, Allele Number, Allele Frequency and Number of Homozygotes as the same of search result. If the variant is also included in dbSNP database, the links of variant are provided. The Browser is linked to the genome browser to present the variant in the region of upstream 100bp and downstream 100bp.

Region annotation

We employed ANNOVAR[1] for the annotation of functionally compiled variants across the entire genome, and utilized the Variant Effect Predictor (VEP)[2] for the annotation of the deleteriousness of variants genome-wide.

Disease annotation

We performed disease annotation of complex variants throughout the entire genome based on linkage disequilibrium (LD), utilizing disease data provided by Genome-Wide Association Studies (GWAS)[3] and the Clinical Variants database (ClinVar)[4].The disease annotation is annotated by clinvar disease database

Pharmacogenomics

The pharmacogenomics variants and related drug information were collected from 34 Clinical Pharmacogenetics Implementation Consortium (CPIC) guidelines (https://cpicpgx.org/). Then add the pharmacogenomics annotation to the variants in this database.

Linkage disequilibrium

R2 and D' statistics between SNP/InDels and TRVs/MEVs/SVs in various populations were calculated and visualized as heatmap plot.

Quantitative trait loci (QTL)

Utilizing 1,085 RNA-seq datasets from individuals in the 1000 Genomes Project, we conducted population cohort study on expression regulation (eQTL) and splicing regulation (sQTL) of structural variations based on the T2T genome.

Population Differentiation

The Fixation index Fst was adopted as the measure for assessing population differences in MEV and SV. Rst was adopted as the measure for assessing population differences in TRV.

The population branch statistics (PBS) was calculated based on Fst statistics, and was adopted as the measure for assessing population differences in genomic region bin around SNP/InDels.

Browser

The browser jumps to genome browser webpage to see the variants coordinates on human genome along with other tracks such as Genes and Gene predictions, Comparative Genomics and variation.

Reference

[1] WANG K, LI M Y, HAKONARSON H. ANNOVAR: functional annotation of genetic variants from high-throughput sequencing data [J]. Nucleic Acids Research, 2010, 38(16).
[2] McLaren W, Gil L, Hunt S E, et al. The ensembl variant effect predictor[J]. Genome biology, 2016, 17: 1-14.
[3] Sollis E, Mosaku A, Abid A, et al. The NHGRI-EBI GWAS Catalog: knowledgebase and deposition resource[J]. Nucleic acids research, 2023, 51(D1): D977-D985.
[4] LANDRUM M J, LEE J M, BENSON M, et al. ClinVar: improving access to variant interpretations and supporting evidence [J]. Nucleic acids research, 2018, 46(D1): D1062-D7.