Novogene Korea
  • Novogene Korea
  • Genomics
    • Human Whole Genome Sequencing
    • Plant & Animal Whole Genome Sequencing
    • Microbial Whole Genome Sequencing
    • Whole Exome Sequencing
    • Plant & Animal De novo Sequencing
    • Microbial De novo Sequencing
    • Amplicon Sequencing
    • Shotgun Metagenomics Sequencing
    Transcriptomics
    • mRNA Sequencing
    • Total RNA Sequencing
    • Full-Length Transcriptome Sequencing
    • Whole Transcriptome Sequencing
    • Small RNA Sequencing
    • Circular RNA Sequencing
    • Metatranscriptome Sequencing
    • Prokaryotic RNA Sequencing
    Single Cell & Spatial Omics
    • Single Cell Gene Expression
    • Single Cell Immune Profiling Sequencing
    • Single Cell Long Read Transcriptome
    • Visium HD Spatial Gene Expression
    • Stereo-Seq Spatial Gene Expression
    • Xenium In Situ Spatial Transcriptome
    Epigenomics
    • Whole Genome Bisulfite Sequencing (WGBS)
    • Directed DNA Methylation Sequencing (DM-Seq) NEW
    • Reduced Representation Bisulfite Sequencing (RRBS)
    • Chromatin Immunoprecipitation Sequencing (ChIP-seq)
    • RNA Immunoprecipitation Sequencing (RIP-seq)
    • Assay for Transposase-Accessible Chromatin with Sequencing (ATAC-seq)

    Premade Library

    • Sequencing Only (Illumina 플랫폼)
    • Sequencing Only (PacBio 플랫폼)
    Proteomics & Metabolomics
    • Olink Proteomics
    • Quantitative Proteomics (MS)
    • Untargeted Metabolomics (MS)
  • 프로모션프로모션
    • 플랫폼
    • 자동화 운송 플랫폼 (Falcon)
    • BI 분석툴 (NovoMagic)
    • Customer Service System (CSS)
    • 브로셔
    • 케이스 스터디
    • 웨비나
    • 블로그
    • 샘플준비 가이드라인
    • 커뮤니티
    • 암 연구
    • 면역 종양학
    • 농업
    • 환경
    • 식품
    • 인간 마이크로바이옴
    • 동물 & 식물 마이크로바이옴
    • 신약개발
    • 희귀 질환 연구
    • 회사소개
    • 글로벌 입지
    • 뉴스룸
    • 채용 정보
  • 문의하기문의하기
  1. Home
  2. Resources
  3. Blog
  4. Tools (GO & KEGG) for Gene Set Enrichment Analysis (GSEA)

Tools (GO & KEGG) for Gene Set Enrichment Analysis (GSEA)

Gene Set Enrichment Analysis (GSEA) is an important tool in genetic research because it can help researchers identify key biological pathways and processes that are associated with a particular phenotype or disease. GSEA is usually employed in genetic research in the following ways:

  • Identifying gene signatures: By analyzing gene expression data using GSEA, researchers can identify gene signatures that are associated with specific phenotypes or diseases. These gene signatures can then be used as diagnostic or prognostic markers, or as potential targets for therapeutic interventions.
  • Understanding disease mechanisms: GSEA can be used to identify biological pathways and processes that are dysregulated in a particular disease or phenotype. This information can help researchers understand the underlying mechanisms of the disease, and can lead to the identification of new therapeutic targets.
  • Drug discovery: GSEA can be used to identify compounds or drugs that are likely to be effective in treating a particular disease or phenotype. By analyzing the gene expression profiles of cells treated with different drugs, researchers can identify drugs that target specific biological pathways or processes that are dysregulated in the disease.

Firstly, the statistical methods commonly used in enrichment analysis include cumulative hypergeometric distribution, Fisher’s exact test, etc. Since a large number of tests (multiple tests) are usually performed simultaneously in enrichment analysis, the test results need to be corrected using multiple test correction methods to make the results more accurate. These methods include Bonferroni correction to counteract the multiple comparisons problem and Benjamini-Hochberg Procedure for false discovery rate correction. The use of enrichment analysis methods to do bioinformatics research on gene annotation databases has generated many enrichment analysis tools, such as DAVID online analysis tool, R Cluster-Profiler package, Meta-scape, etc. These tools play an important role in facilitating the analysis of gene function and the study of biological knowledge data generated by high-throughput sequencing technologies.

The most common GSEA methods currently used are based on enrichment analysis of Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG). Firstly, various techniques are used to multiply a large number of genes of interest, such as differentially expressed gene sets, gene co-expression networks, protein complex gene clusters, etc. In the next step the nodes in GO or pathways of KEGG that are significantly enriched by these gene sets of interest are searched for. This helps in further in-depth and detailed experimental studies. In summary, enrichment analysis is used to decipher the biological knowledge expressed in a set of genes and reveal their roles inside or outside the cell.

Gene Ontology (GO):

The gene ontology (GO) database is a structured standard biological model built by the GO organization in 2000. This model describes our knowledge of biological domains in three aspects that are cellular components, molecular functions and biological processes. It is one of the most widely used gene annotation systems. Each node in the annotation system is a description of a gene or protein. A strict “parent-child” relationship is maintained between the nodes. Thus, a gene or protein can be annotated from three levels.

Fig 1: GO flowchart

  • MF: Molecular Function The molecular activities carried out by gene products. This level focuses on the actions performed rather than the entities that are performing the action.
  • CC: Cellular Component The location where gene products are active and perform a function. This aspect of GO focuses on cellular anatomy rather than processes.
  • BP: Biological Process The larger processes accomplished by the activities of multiple gene products like DNA repair etc. It is to be noted that pathway is not equivalent to biological process. GO does not try to represent the dynamics or dependencies that would be required to fully describe a pathway.
Kyoto Encyclopedia of Genes and Genomes (KEGG)

KEGG is a database for systematic analysis of gene function and genomic information, integrating genomic, biochemical, and phylogenetic information. KEGG is used to understand high-level functions and utilities of the biological system. This database helps researchers study gene and expression of information as a whole. At present, KEGG contains 19 sub-databases. Enrichment analysis is commonly used in KEGG Pathway (It is a collection of manually drawn pathway maps that represent knowledge of the molecular interaction, reaction and relation network). These pathways cover a wide range of biochemical processes.

Fig 2: KEGG flowchart

In conclusion, GO and KEGG are the types of GSEA that are the most frequently used for functional analysis. They are typically the first choice because of their long-standing curation and availability for a wide range of species. They can all be processed through Novomagic’s online tools with just a click.

서비스서비스 menu

고객지원고객지원 menu

기업정보기업정보 menu

서비스
WGSDe novo SeqAmplicon SeqShotgun MetagenomeDM-SeqmRNA-SeqSingle Cell Gene ExpressionVisium HDXenium In SituOlinkUntargeted Metabolomics
고객지원
노보매직CSSFalcon 플랫폼
기업정보
회사소개글로벌 입지플랫폼뉴스룸채용 정보문의하기
LinkedInLinkedIn hoverYouTubeYouTube hoverXX hoverMetaMeta hoverInstagramInstagram hover
Copyright © 2026 Novogene Co., Ltd. All Rights Reserved. 노보진의 한국 내 모든 서비스는 연구 목적 (Research Use Only, RUO) 으로만 제공됩니다. 사업자등록번호: 494-86-03792 | 판매자번호: 노보진코리아유한회사 | 대표자명: 리휘시앙 | 사업자주소: 서울시 강서구 마곡동 779-1번지 뉴브클라우드힐스 BT-230, 231호, 07790 | 전화번호: 02-2038-8036
Novogene Korea
  • Novogene Korea
  • Genomics
    • Human Whole Genome Sequencing
    • Plant & Animal Whole Genome Sequencing
    • Microbial Whole Genome Sequencing
    • Whole Exome Sequencing
    • Plant & Animal De novo Sequencing
    • Microbial De novo Sequencing
    • Amplicon Sequencing
    • Shotgun Metagenomics Sequencing
    Transcriptomics
    • mRNA Sequencing
    • Total RNA Sequencing
    • Full-Length Transcriptome Sequencing
    • Whole Transcriptome Sequencing
    • Small RNA Sequencing
    • Circular RNA Sequencing
    • Metatranscriptome Sequencing
    • Prokaryotic RNA Sequencing
    Single Cell & Spatial Omics
    • Single Cell Gene Expression
    • Single Cell Immune Profiling Sequencing
    • Single Cell Long Read Transcriptome
    • Visium HD Spatial Gene Expression
    • Stereo-Seq Spatial Gene Expression
    • Xenium In Situ Spatial Transcriptome
    Epigenomics
    • Whole Genome Bisulfite Sequencing (WGBS)
    • Directed DNA Methylation Sequencing (DM-Seq) NEW
    • Reduced Representation Bisulfite Sequencing (RRBS)
    • Chromatin Immunoprecipitation Sequencing (ChIP-seq)
    • RNA Immunoprecipitation Sequencing (RIP-seq)
    • Assay for Transposase-Accessible Chromatin with Sequencing (ATAC-seq)

    Premade Library

    • Sequencing Only (Illumina 플랫폼)
    • Sequencing Only (PacBio 플랫폼)
    Proteomics & Metabolomics
    • Olink Proteomics
    • Quantitative Proteomics (MS)
    • Untargeted Metabolomics (MS)
  • 프로모션프로모션
    • 플랫폼
    • 자동화 운송 플랫폼 (Falcon)
    • BI 분석툴 (NovoMagic)
    • Customer Service System (CSS)
    • 브로셔
    • 케이스 스터디
    • 웨비나
    • 블로그
    • 샘플준비 가이드라인
    • 커뮤니티
    • 암 연구
    • 면역 종양학
    • 농업
    • 환경
    • 식품
    • 인간 마이크로바이옴
    • 동물 & 식물 마이크로바이옴
    • 신약개발
    • 희귀 질환 연구
    • 회사소개
    • 글로벌 입지
    • 뉴스룸
    • 채용 정보
  • 문의하기문의하기
  1. Home
  2. Resources
  3. Blog
  4. Tools (GO & KEGG) for Gene Set Enrichment Analysis (GSEA)

Tools (GO & KEGG) for Gene Set Enrichment Analysis (GSEA)

Gene Set Enrichment Analysis (GSEA) is an important tool in genetic research because it can help researchers identify key biological pathways and processes that are associated with a particular phenotype or disease. GSEA is usually employed in genetic research in the following ways:

  • Identifying gene signatures: By analyzing gene expression data using GSEA, researchers can identify gene signatures that are associated with specific phenotypes or diseases. These gene signatures can then be used as diagnostic or prognostic markers, or as potential targets for therapeutic interventions.
  • Understanding disease mechanisms: GSEA can be used to identify biological pathways and processes that are dysregulated in a particular disease or phenotype. This information can help researchers understand the underlying mechanisms of the disease, and can lead to the identification of new therapeutic targets.
  • Drug discovery: GSEA can be used to identify compounds or drugs that are likely to be effective in treating a particular disease or phenotype. By analyzing the gene expression profiles of cells treated with different drugs, researchers can identify drugs that target specific biological pathways or processes that are dysregulated in the disease.

Firstly, the statistical methods commonly used in enrichment analysis include cumulative hypergeometric distribution, Fisher’s exact test, etc. Since a large number of tests (multiple tests) are usually performed simultaneously in enrichment analysis, the test results need to be corrected using multiple test correction methods to make the results more accurate. These methods include Bonferroni correction to counteract the multiple comparisons problem and Benjamini-Hochberg Procedure for false discovery rate correction. The use of enrichment analysis methods to do bioinformatics research on gene annotation databases has generated many enrichment analysis tools, such as DAVID online analysis tool, R Cluster-Profiler package, Meta-scape, etc. These tools play an important role in facilitating the analysis of gene function and the study of biological knowledge data generated by high-throughput sequencing technologies.

The most common GSEA methods currently used are based on enrichment analysis of Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG). Firstly, various techniques are used to multiply a large number of genes of interest, such as differentially expressed gene sets, gene co-expression networks, protein complex gene clusters, etc. In the next step the nodes in GO or pathways of KEGG that are significantly enriched by these gene sets of interest are searched for. This helps in further in-depth and detailed experimental studies. In summary, enrichment analysis is used to decipher the biological knowledge expressed in a set of genes and reveal their roles inside or outside the cell.

Gene Ontology (GO):

The gene ontology (GO) database is a structured standard biological model built by the GO organization in 2000. This model describes our knowledge of biological domains in three aspects that are cellular components, molecular functions and biological processes. It is one of the most widely used gene annotation systems. Each node in the annotation system is a description of a gene or protein. A strict “parent-child” relationship is maintained between the nodes. Thus, a gene or protein can be annotated from three levels.

Fig 1: GO flowchart

  • MF: Molecular Function The molecular activities carried out by gene products. This level focuses on the actions performed rather than the entities that are performing the action.
  • CC: Cellular Component The location where gene products are active and perform a function. This aspect of GO focuses on cellular anatomy rather than processes.
  • BP: Biological Process The larger processes accomplished by the activities of multiple gene products like DNA repair etc. It is to be noted that pathway is not equivalent to biological process. GO does not try to represent the dynamics or dependencies that would be required to fully describe a pathway.
Kyoto Encyclopedia of Genes and Genomes (KEGG)

KEGG is a database for systematic analysis of gene function and genomic information, integrating genomic, biochemical, and phylogenetic information. KEGG is used to understand high-level functions and utilities of the biological system. This database helps researchers study gene and expression of information as a whole. At present, KEGG contains 19 sub-databases. Enrichment analysis is commonly used in KEGG Pathway (It is a collection of manually drawn pathway maps that represent knowledge of the molecular interaction, reaction and relation network). These pathways cover a wide range of biochemical processes.

Fig 2: KEGG flowchart

In conclusion, GO and KEGG are the types of GSEA that are the most frequently used for functional analysis. They are typically the first choice because of their long-standing curation and availability for a wide range of species. They can all be processed through Novomagic’s online tools with just a click.

서비스서비스 menu

고객지원고객지원 menu

기업정보기업정보 menu

서비스
WGSDe novo SeqAmplicon SeqShotgun MetagenomeDM-SeqmRNA-SeqSingle Cell Gene ExpressionVisium HDXenium In SituOlinkUntargeted Metabolomics
고객지원
노보매직CSSFalcon 플랫폼
기업정보
회사소개글로벌 입지플랫폼뉴스룸채용 정보문의하기
LinkedInLinkedIn hoverYouTubeYouTube hoverXX hoverMetaMeta hoverInstagramInstagram hover
Copyright © 2026 Novogene Co., Ltd. All Rights Reserved. 노보진의 한국 내 모든 서비스는 연구 목적 (Research Use Only, RUO) 으로만 제공됩니다. 사업자등록번호: 494-86-03792 | 판매자번호: 노보진코리아유한회사 | 대표자명: 리휘시앙 | 사업자주소: 서울시 강서구 마곡동 779-1번지 뉴브클라우드힐스 BT-230, 231호, 07790 | 전화번호: 02-2038-8036
Privacy PolicyCookie Policy
Privacy PolicyCookie Policy