Welcome to our NyuWa_10K_T2T imputation service!

Introduction

NyuWa, or NüWa, is the mother goddess who was the creator of the human population in Chinese mythology. Here we present NyuWa_10K_T2T haplotype panel, which is the largest and most comprehensive reference panel currently built on the T2T genome. Telomere-to-Telomere (T2T) human genome and 30X sequencing depth enable high quality variant calling, rare variant discovery and SV identification in a fully resolved reference. This resource enables us to explore the population structure and facilitate further genetic studies in populations around the world.

Quick start: Run job, Get results.

Dataset

This version of our NyuWa_10K_T2T haplotype panel was constructed on genomic data of 18,832 haplotypes (9,416 unrelated individuals), which were collected from various projects including 1KGP, SGDP, and HGDP. NyuWa_10K_T2T panel is able to accurately impute variants in T2T-specific regions and structural variants, areas that have historically posed challenges in genome-wide studies. The samples in this resource were mainly sequenced at the target depth of 30X.

Service

Here we present a free web service for users to impute genotype data, utilizing NyuWa resource as reference panel. It provides access to a high-performance computing clusters and to reference panel (NyuWa) of sequenced genomes. You could upload GWS data in VCF format and receive imputed results after processing. Also, we provides a database storing AC, AF, AN and other basic information of variant sites in our NyuWa resource. Users could query variant sites information you need. These are some functions and applications we offer:

  • Mannual page - The instructoin of our website, including how to register and sign in, upload your files(with correct format), get imputed results and use database to query variant sites information. The data security also be stated here.
  • Information querying - Users could query AC, AF, AN and other information of variant sites you are interested in. The basic information will be exhibited in detail after you input position range in genome.
  • Imputation service - You can upload your GWS genotype data directly in this page just following our instruction with VCF format (T2T-CHM13). It should be nocticed that only you have signed in you could upload data and get results. The status of signing in will be kept for one day.

Statement

The imputation service and other applications are totally free, just for scholar communication. In addition, we ensure that the genotype data is only stored for imputation, and we don't retain it for any other purpose. After you have uploaded your VCF files and acquired your imputed results, we will delete your files including your result files in a short time. Of course, you need to register just with simple personal information and sign in before starting jobs in our refpanel for data safety. It means that anyone except you couldn't check your files, imputing status and results. Also, we don't collect your personal information in any shape or form. If you have any questions, please contact us.

Search information

We offer a database storing AC, AF, AN and other basic information of variant sites in our Nyuwa resource. You could enter our page of search directly and conduct querying operation without signing in. You just input the range of sequence in genome, and then we will show the results of variant information in this range. Please note that the range can not be greater than 512 Mb.

Register and sign in

For data safety, you need to register and sign in before you start phasing or imputation jobs in our web server. The information for registeration is streamlined, just for us to authenticate. The state of signing in will last for 24 hours. If you forget your password or have any questions, please contact bigdata@ibp.ac.cn

Start your job

We provide phasing and imputation services for chr1-22 of human genome T2T-CHM13.

You could start your job after you have registered and signed in. We offer three pipelines for your data:

  • phase with Eagle and impute with Beagle5.4
  • phase with Eagle, no imputation
  • impute with Beagle5.4, no phasing

The input file for these pipelines need to be hg38/b38 version vcf/bcf format. It should be noted that the files for imputation only must have been phased, or the data will not be processed.

VCF is a text file format (most likely stored in a compressed manner). It contains meta-information lines (prefixed with "##"), a header line (prefixed with "#") and then data lines each containing information about a position in the genome and genotype information on samples for each position (text fields separated by tabs). Zero length fields are not allowed, a dot (".") must be used instead.

Phase/impute process

Three pipelines are provided for data process with Eagle and Beagle5.4 applying our reference panel.
The Eagle software estimates haplotype phase either within a genotyped cohort or using a phased reference panel. Eagle2 uses a new, very fast HMM-based algorithm that improves speed and accuracy over existing methods via two key ideas: a new data structure based on the positional Burrows-Wheeler transform and a rapid search algorithm that explores only the most relevant paths through the HMM. Compared to the Eagle1 algorithm, Eagle2 has similar speed but much greater accuracy at sample sizes <50,000. Eagle v2.3+ supports phasing sequence data with or without a reference and also supports phasing chrX.
Beagle is a software package for phasing genotypes and imputing ungenotyped markers. Beagle version 5.4 has improved memory and computational efficiency when analyzing large sequence data sets. Beagle supports imputation of multi-allelic variants like TRVs in our panel.

Reference panels

We offer imputation and phasing from these reference panels currently:

  • NyuWa 10k T2T

Get Results

The format of the returned data will be in the Variant Call Format(VCF). Our treatment process begins one minute after the file is uploaded. You could query the status of jobs in Results page and check the results of every job. The log files and processed data could be downloaded. The page will demonstrate your jobs as follows:

After you have click your job, you could check the log files and download the phased/imputated data:

Latest news

  • NyuWaT2T release Oct. 15, 2024
  • The NyuWa resource of 10,000 samples supplies imputation service and information querying of variant sites.