Novogene Korea
  • Novogene Korea
  • Genomics
    • Human Whole Genome Sequencing
    • Plant & Animal Whole Genome Sequencing
    • Microbial Whole Genome Sequencing
    • Whole Exome Sequencing
    • Plant & Animal De novo Sequencing
    • Microbial De novo Sequencing
    • Amplicon Sequencing
    • Shotgun Metagenomics Sequencing
    Transcriptomics
    • mRNA Sequencing
    • Total RNA Sequencing
    • Full-Length Transcriptome Sequencing
    • Whole Transcriptome Sequencing
    • Small RNA Sequencing
    • Circular RNA Sequencing
    • Metatranscriptome Sequencing
    • Prokaryotic RNA Sequencing
    Single Cell & Spatial Omics
    • Single Cell Gene Expression
    • Single Cell Immune Profiling Sequencing
    • Single Cell Long Read Transcriptome
    • Visium HD Spatial Gene Expression
    • Stereo-Seq Spatial Gene Expression
    • Xenium In Situ Spatial Transcriptome
    Epigenomics
    • Whole Genome Bisulfite Sequencing (WGBS)
    • Directed DNA Methylation Sequencing (DM-Seq) NEW
    • Reduced Representation Bisulfite Sequencing (RRBS)
    • Chromatin Immunoprecipitation Sequencing (ChIP-seq)
    • RNA Immunoprecipitation Sequencing (RIP-seq)
    • Assay for Transposase-Accessible Chromatin with Sequencing (ATAC-seq)

    Premade Library

    • Sequencing Only (Illumina 플랫폼)
    • Sequencing Only (PacBio 플랫폼)
    Proteomics & Metabolomics
    • Olink Proteomics
    • Quantitative Proteomics (MS)
    • Untargeted Metabolomics (MS)
  • 프로모션프로모션
    • 플랫폼
    • 자동화 운송 플랫폼 (Falcon)
    • BI 분석툴 (NovoMagic)
    • Customer Service System (CSS)
    • 브로셔
    • 케이스 스터디
    • 웨비나
    • 블로그
    • 샘플준비 가이드라인
    • 커뮤니티
    • 암 연구
    • 면역 종양학
    • 농업
    • 환경
    • 식품
    • 인간 마이크로바이옴
    • 동물 & 식물 마이크로바이옴
    • 신약개발
    • 희귀 질환 연구
    • 회사소개
    • 글로벌 입지
    • 뉴스룸
    • 채용 정보
  • 문의하기문의하기
  1. Home
  2. Resources
  3. Blog
  4. How to choose Normalization methods (TPM/RPKM/FPKM) for mRNA expression

How to choose Normalization methods (TPM/RPKM/FPKM) for mRNA expression

Why do mRNA expression values need to be normalized?

The unification of mRNA expression value measurements across studies, or the normalization of mRNA data, is a significant problem in biomedical and life science research. The abundance of transcripts is measured digitally by reading count. To eliminate technical biases in sequenced data, such as sequencing depth(deeper sequencing depth produces more read counts for one gene) and gene length(longer gene length produces more read counts at the same sequencing level), normalization of gene expression measurements is required.

Notes:● READS COUNTS: Obtained from the original sequencing data, the count number is the total number of reads mapped to a certain gene; in the sequencing analysis process, the measured short reads are firstly mapped to the reference genome, and then the software is used to calculate the number of reads mapped to a certain gene, which means that read count is an integer value.

RPKM (Reads Per Kilobase per Million mapped reads)was made for single-end RNA-seq, where every read corresponded to a single fragment that was sequenced. FPKM (Fragments Per Kilobase per Million mapped fragments) is very similar to RPKM. We divide the number of fragments of a gene by the total sequencing depth, and the ratio is divided by the gene length. Note that, strictly speaking, the gene length mentioned above represents the total length of exons from one gene.

The difference between RPKM and FPKM is that F stands for fragments and R stands for reads. In the case of PE (Pair-end) sequencing, each fragment will have two reads, and FPKM only calculates the number of fragments that can be compared to the same transcript for both reads, while RPKM calculates the number of reads that can be compared to the transcript. The FPKM only counts the number of fragments that can be matched to the same transcript. In the case of SE (single-end) sequencing, the results calculated by FPKM and RPKM will be the same.

FPKM and RPKM ultimately normalize the abundance of transcripts from different samples (or the same sample under different conditions) to a standard that allows quantitative comparison by dividing both L (transcript length) and N (total number of Reads (Fragment)).

TPM (transcripts per kilobase million) is very much like FPKM and RPKM, but the only difference is that at first, normalize for gene length, and later normalize for sequencing depth. However, the differencing effect is very profound. Therefore, TPM is a more accurate statistic when calculating gene expression comparisons across samples. While using TPM, the sum of all TPMs are the same in each sample. This makes the comparison of the proportion of reads mapped to a gene in each sample very convenient.

How to choose the normalization method?

The TPM normalization results are sample independent and the TPMs are guaranteed to be the same across samples; however, the FPKM and TPM are about the same for each gene in each sample, so many people still use FPKM or RPKM to compare expression values of the same gene across samples. As with any high sequencing throughput technology, the analytical method is critical to interpret the data, and the RNA-seq analysis process is always evolving. Therefore, the appropriate method should be selected based on a combination of research directions.

Normalization methodDescriptionRecommendations for use
TPM (transcripts per kilobase million)Counts per length of transcript (kb) per million reads mappedGene count comparisons within a sample or between samples of the same sample group;
RPKM/FPKM (reads/fragments per kilobase per million reads/fragments mapped)Normalize for gene length at first, and later normalize for sequencing depthGene count comparisons between genes within as ample; NOT for between sample comparisons

Reference

  1. Dillies, Marie-Agnès, et al. “A comprehensive evaluation of normalization methods for Illumina high-throughput RNA sequencing data analysis.” Briefings in bioinformatics 14.6 (2013): 671-683.
  2. 2.Fundel, K., et al. “Normalization strategies for mRNA expression data in cartilage research.” Osteoarthritis and cartilage 16.8 (2008): 947-955.

서비스서비스 menu

고객지원고객지원 menu

기업정보기업정보 menu

서비스
WGSDe novo SeqAmplicon SeqShotgun MetagenomeDM-SeqmRNA-SeqSingle Cell Gene ExpressionVisium HDXenium In SituOlinkUntargeted Metabolomics
고객지원
노보매직CSSFalcon 플랫폼
기업정보
회사소개글로벌 입지플랫폼뉴스룸채용 정보문의하기
LinkedInLinkedIn hoverYouTubeYouTube hoverXX hoverMetaMeta hoverInstagramInstagram hover
Copyright © 2026 Novogene Co., Ltd. All Rights Reserved. 노보진의 한국 내 모든 서비스는 연구 목적 (Research Use Only, RUO) 으로만 제공됩니다. 사업자등록번호: 494-86-03792 | 판매자번호: 노보진코리아유한회사 | 대표자명: 리휘시앙 | 사업자주소: 서울시 강서구 마곡동 779-1번지 뉴브클라우드힐스 BT-230, 231호, 07790 | 전화번호: 02-2038-8036
Novogene Korea
  • Novogene Korea
  • Genomics
    • Human Whole Genome Sequencing
    • Plant & Animal Whole Genome Sequencing
    • Microbial Whole Genome Sequencing
    • Whole Exome Sequencing
    • Plant & Animal De novo Sequencing
    • Microbial De novo Sequencing
    • Amplicon Sequencing
    • Shotgun Metagenomics Sequencing
    Transcriptomics
    • mRNA Sequencing
    • Total RNA Sequencing
    • Full-Length Transcriptome Sequencing
    • Whole Transcriptome Sequencing
    • Small RNA Sequencing
    • Circular RNA Sequencing
    • Metatranscriptome Sequencing
    • Prokaryotic RNA Sequencing
    Single Cell & Spatial Omics
    • Single Cell Gene Expression
    • Single Cell Immune Profiling Sequencing
    • Single Cell Long Read Transcriptome
    • Visium HD Spatial Gene Expression
    • Stereo-Seq Spatial Gene Expression
    • Xenium In Situ Spatial Transcriptome
    Epigenomics
    • Whole Genome Bisulfite Sequencing (WGBS)
    • Directed DNA Methylation Sequencing (DM-Seq) NEW
    • Reduced Representation Bisulfite Sequencing (RRBS)
    • Chromatin Immunoprecipitation Sequencing (ChIP-seq)
    • RNA Immunoprecipitation Sequencing (RIP-seq)
    • Assay for Transposase-Accessible Chromatin with Sequencing (ATAC-seq)

    Premade Library

    • Sequencing Only (Illumina 플랫폼)
    • Sequencing Only (PacBio 플랫폼)
    Proteomics & Metabolomics
    • Olink Proteomics
    • Quantitative Proteomics (MS)
    • Untargeted Metabolomics (MS)
  • 프로모션프로모션
    • 플랫폼
    • 자동화 운송 플랫폼 (Falcon)
    • BI 분석툴 (NovoMagic)
    • Customer Service System (CSS)
    • 브로셔
    • 케이스 스터디
    • 웨비나
    • 블로그
    • 샘플준비 가이드라인
    • 커뮤니티
    • 암 연구
    • 면역 종양학
    • 농업
    • 환경
    • 식품
    • 인간 마이크로바이옴
    • 동물 & 식물 마이크로바이옴
    • 신약개발
    • 희귀 질환 연구
    • 회사소개
    • 글로벌 입지
    • 뉴스룸
    • 채용 정보
  • 문의하기문의하기
  1. Home
  2. Resources
  3. Blog
  4. How to choose Normalization methods (TPM/RPKM/FPKM) for mRNA expression

How to choose Normalization methods (TPM/RPKM/FPKM) for mRNA expression

Why do mRNA expression values need to be normalized?

The unification of mRNA expression value measurements across studies, or the normalization of mRNA data, is a significant problem in biomedical and life science research. The abundance of transcripts is measured digitally by reading count. To eliminate technical biases in sequenced data, such as sequencing depth(deeper sequencing depth produces more read counts for one gene) and gene length(longer gene length produces more read counts at the same sequencing level), normalization of gene expression measurements is required.

Notes:● READS COUNTS: Obtained from the original sequencing data, the count number is the total number of reads mapped to a certain gene; in the sequencing analysis process, the measured short reads are firstly mapped to the reference genome, and then the software is used to calculate the number of reads mapped to a certain gene, which means that read count is an integer value.

RPKM (Reads Per Kilobase per Million mapped reads)was made for single-end RNA-seq, where every read corresponded to a single fragment that was sequenced. FPKM (Fragments Per Kilobase per Million mapped fragments) is very similar to RPKM. We divide the number of fragments of a gene by the total sequencing depth, and the ratio is divided by the gene length. Note that, strictly speaking, the gene length mentioned above represents the total length of exons from one gene.

The difference between RPKM and FPKM is that F stands for fragments and R stands for reads. In the case of PE (Pair-end) sequencing, each fragment will have two reads, and FPKM only calculates the number of fragments that can be compared to the same transcript for both reads, while RPKM calculates the number of reads that can be compared to the transcript. The FPKM only counts the number of fragments that can be matched to the same transcript. In the case of SE (single-end) sequencing, the results calculated by FPKM and RPKM will be the same.

FPKM and RPKM ultimately normalize the abundance of transcripts from different samples (or the same sample under different conditions) to a standard that allows quantitative comparison by dividing both L (transcript length) and N (total number of Reads (Fragment)).

TPM (transcripts per kilobase million) is very much like FPKM and RPKM, but the only difference is that at first, normalize for gene length, and later normalize for sequencing depth. However, the differencing effect is very profound. Therefore, TPM is a more accurate statistic when calculating gene expression comparisons across samples. While using TPM, the sum of all TPMs are the same in each sample. This makes the comparison of the proportion of reads mapped to a gene in each sample very convenient.

How to choose the normalization method?

The TPM normalization results are sample independent and the TPMs are guaranteed to be the same across samples; however, the FPKM and TPM are about the same for each gene in each sample, so many people still use FPKM or RPKM to compare expression values of the same gene across samples. As with any high sequencing throughput technology, the analytical method is critical to interpret the data, and the RNA-seq analysis process is always evolving. Therefore, the appropriate method should be selected based on a combination of research directions.

Normalization methodDescriptionRecommendations for use
TPM (transcripts per kilobase million)Counts per length of transcript (kb) per million reads mappedGene count comparisons within a sample or between samples of the same sample group;
RPKM/FPKM (reads/fragments per kilobase per million reads/fragments mapped)Normalize for gene length at first, and later normalize for sequencing depthGene count comparisons between genes within as ample; NOT for between sample comparisons

Reference

  1. Dillies, Marie-Agnès, et al. “A comprehensive evaluation of normalization methods for Illumina high-throughput RNA sequencing data analysis.” Briefings in bioinformatics 14.6 (2013): 671-683.
  2. 2.Fundel, K., et al. “Normalization strategies for mRNA expression data in cartilage research.” Osteoarthritis and cartilage 16.8 (2008): 947-955.

서비스서비스 menu

고객지원고객지원 menu

기업정보기업정보 menu

서비스
WGSDe novo SeqAmplicon SeqShotgun MetagenomeDM-SeqmRNA-SeqSingle Cell Gene ExpressionVisium HDXenium In SituOlinkUntargeted Metabolomics
고객지원
노보매직CSSFalcon 플랫폼
기업정보
회사소개글로벌 입지플랫폼뉴스룸채용 정보문의하기
LinkedInLinkedIn hoverYouTubeYouTube hoverXX hoverMetaMeta hoverInstagramInstagram hover
Copyright © 2026 Novogene Co., Ltd. All Rights Reserved. 노보진의 한국 내 모든 서비스는 연구 목적 (Research Use Only, RUO) 으로만 제공됩니다. 사업자등록번호: 494-86-03792 | 판매자번호: 노보진코리아유한회사 | 대표자명: 리휘시앙 | 사업자주소: 서울시 강서구 마곡동 779-1번지 뉴브클라우드힐스 BT-230, 231호, 07790 | 전화번호: 02-2038-8036
Privacy PolicyCookie Policy
Privacy PolicyCookie Policy