Novogene Korea
  • Novogene Korea
  • Genomics
    • Human Whole Genome Sequencing
    • Plant & Animal Whole Genome Sequencing
    • Microbial Whole Genome Sequencing
    • Whole Exome Sequencing
    • Plant & Animal De novo Sequencing
    • Microbial De novo Sequencing
    • Amplicon Sequencing
    • Shotgun Metagenomics Sequencing
    Transcriptomics
    • mRNA Sequencing
    • Total RNA Sequencing
    • Full-Length Transcriptome Sequencing
    • Whole Transcriptome Sequencing
    • Small RNA Sequencing
    • Circular RNA Sequencing
    • Metatranscriptome Sequencing
    • Prokaryotic RNA Sequencing
    Single Cell & Spatial Omics
    • Single Cell Gene Expression
    • Single Cell Immune Profiling Sequencing
    • Single Cell Long Read Transcriptome
    • Visium HD Spatial Gene Expression
    • Stereo-Seq Spatial Gene Expression
    • Xenium In Situ Spatial Transcriptome
    Epigenomics
    • Whole Genome Bisulfite Sequencing (WGBS)
    • Directed DNA Methylation Sequencing (DM-Seq) NEW
    • Reduced Representation Bisulfite Sequencing (RRBS)
    • Chromatin Immunoprecipitation Sequencing (ChIP-seq)
    • RNA Immunoprecipitation Sequencing (RIP-seq)
    • Assay for Transposase-Accessible Chromatin with Sequencing (ATAC-seq)

    Premade Library

    • Sequencing Only (Illumina 플랫폼)
    • Sequencing Only (PacBio 플랫폼)
    Proteomics & Metabolomics
    • Olink Proteomics
    • Quantitative Proteomics (MS)
    • Untargeted Metabolomics (MS)
  • 프로모션프로모션
    • 플랫폼
    • 자동화 운송 플랫폼 (Falcon)
    • BI 분석툴 (NovoMagic)
    • Customer Service System (CSS)
    • 브로셔
    • 케이스 스터디
    • 웨비나
    • 블로그
    • 샘플준비 가이드라인
    • 커뮤니티
    • 암 연구
    • 면역 종양학
    • 농업
    • 환경
    • 식품
    • 인간 마이크로바이옴
    • 동물 & 식물 마이크로바이옴
    • 신약개발
    • 희귀 질환 연구
    • 회사소개
    • 글로벌 입지
    • 뉴스룸
    • 채용 정보
  • 문의하기문의하기
  1. Home
  2. Resources
  3. Blog
  4. A Beginners Guide to DNA-seq: Bioinformatics Analysis

A Beginners Guide to DNA-seq: Bioinformatics Analysis

Next-generation sequencing (NGS) technologies have advanced scientific experiments by enabling large-scale genomic sequencing projects to be carried out. The study of genetic variation is applicable across a wide range of fields and can be used to answer questions from ecology and evolution to examining various diseases.

DNA sequencing at Novogene

DNA sequencing is used to determine the sequence of the four bases that make up the DNA molecule. The advancement of next-generation sequencing has made it a popular tool for answering a variety of questions in the life science and medical fields.

Novogene offers a range of DNA sequencing services on three different platforms. The Illumina platform is used for short reads and for long reads we use both the PacBio and ONT systems. The type of sequencing platform used will depend on the type of data that you are sequencing and will need to take into consideration how much data is required to determine the breadth and depth of the sequencing reads. At Novogene, we offer both whole genome and whole exome sequencing. Whole genome sequencing will provide you with data on all the genes as well as the non-coding regions, while the whole exome only examines the DNA contained in exomes and can be more cost-effective depending on the type of data you are looking for.

An overview of DNA-Seq Bioinformatics Analysis: Germline vs Somatic

At Novogene, we can also perform the bioinformatics analysis on your samples. To give you an idea of what that analysis will look like let’s briefly go over some of the pipelines that we use. The first variant calling workflow we will examine is the Germline Variant Calling workflow and this is used when carrying out short reads on the Illumina platform. This pipeline consists of five steps:

1.Data QC

2. Alignment

3.Remove Duplicates

4.SNP/InDEL calling

5.Annotation.

In the first step, we clean up the raw reads before using a program called BWA to carry out an alignment. This can be either a local alignment or global alignment depending on the type of reads produced. Once aligned, the files are converted to a BAM file and assessed for duplicates. Finally, we check the quality of the data by calculating a base quality score recalibration. Once this has been done, we use a program called GATK4 haplotype caller which is one of the common tools for germline variant calling. This uses a sliding window across the reference genome to identify the active region. This generates a genomic VCF file that can be used for joint genotyping. In addition, from this, we can call variants across all samples to get a final VCF file. Another tool we can use for germline variant calling the Deep Variant. This tool is also used with the BAM or reference file and produces a VCF file. The VCF file can then be annotated with either genome annotation or region-based annotation depending on your needs.

The second type of workflow that we are going to talk you through is somatic variant calling. This is a whole different story to the germline pipeline as we are dealing with low-frequency variants which need to be identified and separated from artifacts. An example of a basic somatic variant calling pipeline is the VarScan2 pipeline. Here we start with BAM files or Germline population resources and perform SNV and InDEL calling. This produces an Indel VCF or an SNV VCF depending on what you start with. Variant filtering can then be used to produce analysis-ready variants.

Long reads and advanced analysis

For longer reads, we use the PacBio and Nanopore variant calling. These are invaluable tools for the discovery of structured variants (SV). Structured variants are variations within the human genome that exceed 50 base pairs. SVs are identified using one of four methods:

1.Read depth

2.Paired reads

3.Split reads

4.De novo assembly

These methods work in different ways depending on the data that you have to identify areas where there have been insertions or deletions in the DNA sequence. In addition to these analyses, we also offer advanced analyses such as Mendelian disease analysis and in-depth cancer analysis.

Novogene is a world expert in the sequencing field and will provide you with a comprehensive service that includes recommendations on sequencing and bioinformatics depending on the types of samples that you wish to process.

For more information on variant calling pipelines for Illumina, PacBio, and Nanopore data and our advanced analysis approaches for disease and cancer studies you can listen to our webinar available here:A Beginner’s Guide to DNA-seq Bioinformatics Analysis – Novogene Feel free to learn more about Whole Genome Sequencing here:Novogene Whole Genome Sequencing

서비스서비스 menu

고객지원고객지원 menu

기업정보기업정보 menu

서비스
WGSDe novo SeqAmplicon SeqShotgun MetagenomeDM-SeqmRNA-SeqSingle Cell Gene ExpressionVisium HDXenium In SituOlinkUntargeted Metabolomics
고객지원
노보매직CSSFalcon 플랫폼
기업정보
회사소개글로벌 입지플랫폼뉴스룸채용 정보문의하기
LinkedInLinkedIn hoverYouTubeYouTube hoverXX hoverMetaMeta hoverInstagramInstagram hover
Copyright © 2026 Novogene Co., Ltd. All Rights Reserved. 노보진의 한국 내 모든 서비스는 연구 목적 (Research Use Only, RUO) 으로만 제공됩니다. 사업자등록번호: 494-86-03792 | 판매자번호: 노보진코리아유한회사 | 대표자명: 리휘시앙 | 사업자주소: 서울시 강서구 마곡동 779-1번지 뉴브클라우드힐스 BT-230, 231호, 07790 | 전화번호: 02-2038-8036
Novogene Korea
  • Novogene Korea
  • Genomics
    • Human Whole Genome Sequencing
    • Plant & Animal Whole Genome Sequencing
    • Microbial Whole Genome Sequencing
    • Whole Exome Sequencing
    • Plant & Animal De novo Sequencing
    • Microbial De novo Sequencing
    • Amplicon Sequencing
    • Shotgun Metagenomics Sequencing
    Transcriptomics
    • mRNA Sequencing
    • Total RNA Sequencing
    • Full-Length Transcriptome Sequencing
    • Whole Transcriptome Sequencing
    • Small RNA Sequencing
    • Circular RNA Sequencing
    • Metatranscriptome Sequencing
    • Prokaryotic RNA Sequencing
    Single Cell & Spatial Omics
    • Single Cell Gene Expression
    • Single Cell Immune Profiling Sequencing
    • Single Cell Long Read Transcriptome
    • Visium HD Spatial Gene Expression
    • Stereo-Seq Spatial Gene Expression
    • Xenium In Situ Spatial Transcriptome
    Epigenomics
    • Whole Genome Bisulfite Sequencing (WGBS)
    • Directed DNA Methylation Sequencing (DM-Seq) NEW
    • Reduced Representation Bisulfite Sequencing (RRBS)
    • Chromatin Immunoprecipitation Sequencing (ChIP-seq)
    • RNA Immunoprecipitation Sequencing (RIP-seq)
    • Assay for Transposase-Accessible Chromatin with Sequencing (ATAC-seq)

    Premade Library

    • Sequencing Only (Illumina 플랫폼)
    • Sequencing Only (PacBio 플랫폼)
    Proteomics & Metabolomics
    • Olink Proteomics
    • Quantitative Proteomics (MS)
    • Untargeted Metabolomics (MS)
  • 프로모션프로모션
    • 플랫폼
    • 자동화 운송 플랫폼 (Falcon)
    • BI 분석툴 (NovoMagic)
    • Customer Service System (CSS)
    • 브로셔
    • 케이스 스터디
    • 웨비나
    • 블로그
    • 샘플준비 가이드라인
    • 커뮤니티
    • 암 연구
    • 면역 종양학
    • 농업
    • 환경
    • 식품
    • 인간 마이크로바이옴
    • 동물 & 식물 마이크로바이옴
    • 신약개발
    • 희귀 질환 연구
    • 회사소개
    • 글로벌 입지
    • 뉴스룸
    • 채용 정보
  • 문의하기문의하기
  1. Home
  2. Resources
  3. Blog
  4. A Beginners Guide to DNA-seq: Bioinformatics Analysis

A Beginners Guide to DNA-seq: Bioinformatics Analysis

Next-generation sequencing (NGS) technologies have advanced scientific experiments by enabling large-scale genomic sequencing projects to be carried out. The study of genetic variation is applicable across a wide range of fields and can be used to answer questions from ecology and evolution to examining various diseases.

DNA sequencing at Novogene

DNA sequencing is used to determine the sequence of the four bases that make up the DNA molecule. The advancement of next-generation sequencing has made it a popular tool for answering a variety of questions in the life science and medical fields.

Novogene offers a range of DNA sequencing services on three different platforms. The Illumina platform is used for short reads and for long reads we use both the PacBio and ONT systems. The type of sequencing platform used will depend on the type of data that you are sequencing and will need to take into consideration how much data is required to determine the breadth and depth of the sequencing reads. At Novogene, we offer both whole genome and whole exome sequencing. Whole genome sequencing will provide you with data on all the genes as well as the non-coding regions, while the whole exome only examines the DNA contained in exomes and can be more cost-effective depending on the type of data you are looking for.

An overview of DNA-Seq Bioinformatics Analysis: Germline vs Somatic

At Novogene, we can also perform the bioinformatics analysis on your samples. To give you an idea of what that analysis will look like let’s briefly go over some of the pipelines that we use. The first variant calling workflow we will examine is the Germline Variant Calling workflow and this is used when carrying out short reads on the Illumina platform. This pipeline consists of five steps:

1.Data QC

2. Alignment

3.Remove Duplicates

4.SNP/InDEL calling

5.Annotation.

In the first step, we clean up the raw reads before using a program called BWA to carry out an alignment. This can be either a local alignment or global alignment depending on the type of reads produced. Once aligned, the files are converted to a BAM file and assessed for duplicates. Finally, we check the quality of the data by calculating a base quality score recalibration. Once this has been done, we use a program called GATK4 haplotype caller which is one of the common tools for germline variant calling. This uses a sliding window across the reference genome to identify the active region. This generates a genomic VCF file that can be used for joint genotyping. In addition, from this, we can call variants across all samples to get a final VCF file. Another tool we can use for germline variant calling the Deep Variant. This tool is also used with the BAM or reference file and produces a VCF file. The VCF file can then be annotated with either genome annotation or region-based annotation depending on your needs.

The second type of workflow that we are going to talk you through is somatic variant calling. This is a whole different story to the germline pipeline as we are dealing with low-frequency variants which need to be identified and separated from artifacts. An example of a basic somatic variant calling pipeline is the VarScan2 pipeline. Here we start with BAM files or Germline population resources and perform SNV and InDEL calling. This produces an Indel VCF or an SNV VCF depending on what you start with. Variant filtering can then be used to produce analysis-ready variants.

Long reads and advanced analysis

For longer reads, we use the PacBio and Nanopore variant calling. These are invaluable tools for the discovery of structured variants (SV). Structured variants are variations within the human genome that exceed 50 base pairs. SVs are identified using one of four methods:

1.Read depth

2.Paired reads

3.Split reads

4.De novo assembly

These methods work in different ways depending on the data that you have to identify areas where there have been insertions or deletions in the DNA sequence. In addition to these analyses, we also offer advanced analyses such as Mendelian disease analysis and in-depth cancer analysis.

Novogene is a world expert in the sequencing field and will provide you with a comprehensive service that includes recommendations on sequencing and bioinformatics depending on the types of samples that you wish to process.

For more information on variant calling pipelines for Illumina, PacBio, and Nanopore data and our advanced analysis approaches for disease and cancer studies you can listen to our webinar available here:A Beginner’s Guide to DNA-seq Bioinformatics Analysis – Novogene Feel free to learn more about Whole Genome Sequencing here:Novogene Whole Genome Sequencing

서비스서비스 menu

고객지원고객지원 menu

기업정보기업정보 menu

서비스
WGSDe novo SeqAmplicon SeqShotgun MetagenomeDM-SeqmRNA-SeqSingle Cell Gene ExpressionVisium HDXenium In SituOlinkUntargeted Metabolomics
고객지원
노보매직CSSFalcon 플랫폼
기업정보
회사소개글로벌 입지플랫폼뉴스룸채용 정보문의하기
LinkedInLinkedIn hoverYouTubeYouTube hoverXX hoverMetaMeta hoverInstagramInstagram hover
Copyright © 2026 Novogene Co., Ltd. All Rights Reserved. 노보진의 한국 내 모든 서비스는 연구 목적 (Research Use Only, RUO) 으로만 제공됩니다. 사업자등록번호: 494-86-03792 | 판매자번호: 노보진코리아유한회사 | 대표자명: 리휘시앙 | 사업자주소: 서울시 강서구 마곡동 779-1번지 뉴브클라우드힐스 BT-230, 231호, 07790 | 전화번호: 02-2038-8036
Privacy PolicyCookie Policy
Privacy PolicyCookie Policy