RNA-seq has become a routine tool in plant research. Researchers compare tissues and developmental stages, expose plants to stress, and contrast mutants with wild type. As a result, enormous amounts of RNA-seq data have accumulated in public repositories such as the SRA.
But a FASTQ file being public is not the same as that dataset being immediately reusable. A researcher still has to determine which tissue was sampled, what treatment was applied, which cultivar was used, retrieve the data, perform quality control, map the reads, count them, and normalize expression values.
On September 2, 2026, an accepted manuscript of a Database Paper describing the Grass Expression Atlas (GExA) was published in Plant and Cell Physiology. GExA is research infrastructure designed to lower this barrier for grasses, with particularly strong coverage of millets.
Its broader significance is not simply that it provides another expression database. It illustrates a shift from only generating new RNA-seq data toward turning the RNA-seq data that already exist around the world into reusable research assets.
- Turning 4,753 public RNA-seq samples into searchable data
- Starting with a simple question: “What might this gene be doing?”
- A first-pass screen before generating another RNA-seq dataset
- 4,753 samples are not 4,753 biological replicates
- It is not a direct cross-species expression-comparison platform
- A reference genome can change what expression looks like
- Reusing public RNA-seq is not a new idea
- Before generating new RNA-seq, look at what already exists
- After “make the data public” comes “make the data usable”
- References and resources
Turning 4,753 public RNA-seq samples into searchable data
The current GExA release contains seven grass species: pearl millet, foxtail millet, proso millet, finger millet, sorghum, barley, and rice. Together, they account for 4,753 samples from 457 BioProjects.
The largest collections are foxtail millet, with 2,216 samples, and pearl millet, with 987. Rice currently contributes 80 samples. Wheat and maize are not included in the present release.
So despite the broad name “Grass Expression Atlas,” its present strength is more specific: it makes public transcriptome data from millets substantially easier to explore.
GExA is not merely a directory of accession numbers. Public RNA-seq reads are processed through a standardized workflow to generate gene-level read counts and TPM values, while metadata such as tissue, treatment, developmental stage, and cultivar are organized for searching and grouping.
In other words, it turns
a collection of SRA accessions
into
data that can answer, at least at an exploratory level, where, when, and under which conditions a gene was expressed.
That transformation is the core of GExA.
Starting with a simple question: “What might this gene be doing?”
A researcher can enter a Gene ID and inspect its expression grouped by tissue, treatment, developmental stage, cultivar, and other metadata. Multiple genes can also be viewed together.
When the Gene ID is unknown, GExA can use annotation keywords derived from BLASTP homology to rice and Arabidopsis proteins to identify candidate genes. Individual samples can then be traced back through their metadata to the original BioProject.
The official Help page uses pearl millet gene dpca1g022930.840 as an example. It is a close homolog of rice HVA22, a gene associated with dehydration and ABA responses. GExA makes it possible to inspect multiple drought- and ABA-related public samples in which this candidate appears.
The point is not to conclude from the atlas alone that “this is a drought-tolerance gene.”
The useful step is to find an interesting condition and then return to the original experiment.
A first-pass screen before generating another RNA-seq dataset
Suppose a newly identified candidate gene looks as though it might be involved in drought tolerance.
Traditionally, a researcher might search the literature, inspect supplementary datasets, and eventually download public RNA-seq data for reanalysis.
With GExA, several questions can be asked first:
- Is the gene expressed in roots?
- Are there drought-related samples with high expression?
- Is it also strongly expressed in reproductive tissues?
- Is the pattern restricted to one cultivar?
If an interesting pattern appears, the researcher can move back to the relevant BioProject and paper, inspect the experimental design, and download the raw data when a rigorous reanalysis is warranted.
The workflow becomes:
Explore in the atlas → inspect the original BioProject and paper → perform a dedicated analysis if needed.
GExA is therefore less a tool for producing final conclusions than a tool for deciding what is worth investigating next.
4,753 samples are not 4,753 biological replicates
This distinction is critical.
The 4,753 samples were not generated in one giant controlled experiment. They come from 457 separate BioProjects.
Laboratories, growth conditions, RNA extraction protocols, library preparation methods, sequencing platforms, read lengths, and experimental designs differ. Reprocessing all datasets with a common workflow and expressing them as TPM does not eliminate these batch effects and confounders.
The authors explicitly acknowledge this limitation. GExA is primarily intended for exploratory analysis and hypothesis generation. Rigorous differential-expression analysis or causal interpretation should be supported by appropriately controlled independent datasets.
For example, comparing samples from different BioProjects and observing a fivefold TPM difference does not justify the statement that “drought induces this gene fivefold.”
What the atlas can support is a softer but useful observation such as:
“High expression appears repeatedly in several drought-related datasets,”
or
“The strongest signals seem concentrated in root-derived samples.”
Those observations can generate a hypothesis.
Exploration, not proof.
That is the appropriate distance to keep when reusing heterogeneous public RNA-seq data.
It is not a direct cross-species expression-comparison platform
Another potential overinterpretation is to describe GExA as a full cross-species expression-comparison system.
Multiple grass species are present in the same resource, and homology annotations to rice and Arabidopsis help researchers identify related genes. But the current interface is not primarily designed to automatically pair orthologs and statistically compare normalized expression levels across species.
Its cross-species value is more modest and still useful: it makes it easier to explore homologous genes and their expression information across multiple grass species.
A rigorous comparative transcriptomics analysis still requires explicit decisions about orthology, tissue equivalence, developmental stage, normalization, and other sources of biological and technical variation.
A reference genome can change what expression looks like
Pearl millet provides another technically interesting feature.
GExA supports three reference genomes for pearl millet: Tift, ICMB843, and ICMR06777.
In RNA-seq analysis, sequence divergence between a sample and the chosen reference can affect mapping efficiency and therefore influence estimated expression. A low value can reflect genuinely low expression, but it can also be affected by poor mapping caused by reference divergence.
GExA allows pearl millet expression to be inspected against multiple references and provides links to high-similarity genes in other cultivars.
Public RNA-seq is therefore not a case where “the same FASTQ always produces the same biological answer.”
The reference genome is part of the analysis.
GExA’s multi-reference design makes that issue unusually visible.
Reusing public RNA-seq is not a new idea
GExA should not be overcredited for inventing this approach.
Projects have been collecting, reprocessing, and making public transcriptomes searchable for years.
EMBL-EBI’s Expression Atlas has long provided access to expression studies across many species.
PlantExp, reported in 2023, standardized 131,423 public RNA-seq samples from 85 plant species and supports expression exploration, differential expression, co-expression, alternative splicing, and cross-species expression conservation.
GExA is therefore neither the first public-RNA-seq reuse platform nor the largest plant expression database.
Its importance lies elsewhere.
This kind of infrastructure is now reaching crops such as millets, where bioinformatics resources have historically been less extensive than in the best-supported model species and major crops.
Rather than competing with the largest general-purpose databases, GExA fills a practical gap: it turns public data from relatively underserved grass species into something researchers can explore quickly.
Before generating new RNA-seq, look at what already exists
None of this means that new RNA-seq experiments are becoming unnecessary.
Testing a specific hypothesis still requires appropriate controls, the right tissue, genotype, developmental stage, sampling time, and experimental design.
What can change is the step before that experiment.
A researcher identifies a candidate gene.
Instead of immediately designing a new RNA-seq experiment, the researcher first checks public data.
Existing datasets are used to narrow the hypothesis and identify missing conditions.
Then new RNA-seq is generated for the question that existing data cannot answer.
That sequence can make newly generated data more valuable, not less.
The shift is therefore not from “new RNA-seq” to “no new RNA-seq.” It is from a research culture focused mainly on producing datasets toward one that first asks how much can be learned from the datasets already available.
Explore and reuse accumulated RNA-seq first, then generate the data that are genuinely needed.
GExA makes that transition tangible in grass research, especially for millets.
After “make the data public” comes “make the data usable”
Depositing FASTQ files in a public repository is now standard practice, but publication alone does not remove the barriers to reuse.
Metadata need to be organized. Data need to be processed consistently. Researchers need to be able to search by Gene ID or annotation and then trace a signal back to the original experiment.
Only then can an RNA-seq dataset generated years ago participate in a new research question.
What GExA has built is therefore not primarily a new RNA-seq dataset. It is a layer between
RNA-seq data that already existed
and
the next research hypothesis.
The same group has also released a 2026 preprint describing DEAR-OWL, a browser-based differential-expression tool that can use GExA data as built-in datasets. It remains a preprint, but the direction is notable: an atlas can evolve from a place where researchers look at public data into an environment where they can begin reanalyzing those data directly.
The value of research data does not end when the original paper is published.
RNA-seq generated years ago by another laboratory may help prioritize a candidate gene, reshape an experimental design, or generate a new hypothesis today.
That requires more than producing ever more data. It requires infrastructure that turns existing data into usable knowledge.
Grass Expression Atlas is not a paper built around a dramatic new biological discovery. But if “check existing public RNA-seq before generating the next dataset” becomes a normal part of plant-research workflows, resources like this may have substantial long-term impact.
References and resources
- Kambara K, Chen Q, Tsugama D. Grass Expression Atlas: an RNA-seq-based expression resource for grass species. Plant and Cell Physiology. 2026. DOI: 10.1093/pcp/pcag122. https://academic.oup.com/pcp/advance-article/doi/10.1093/pcp/pcag122/8780316
- Grass Expression Atlas: https://www.gexa.anesc.u-tokyo.ac.jp/
- GExA About: https://www.gexa.anesc.u-tokyo.ac.jp/About/index.html
- GExA Help: https://www.gexa.anesc.u-tokyo.ac.jp/Help/index.html
- GExA Download: https://www.gexa.anesc.u-tokyo.ac.jp/Download/index.html
- Ma S, et al. PlantExp: a platform for exploration of gene expression and alternative splicing profiles in 85 plant species. Nucleic Acids Research (2023). https://academic.oup.com/nar/article/51/D1/D1483/6769746
- Expression Atlas update: https://pmc.ncbi.nlm.nih.gov/articles/PMC12807774/
- DEAR-OWL preprint: https://www.biorxiv.org/content/10.64898/2026.07.28.741369v1


Comments