What Is PeptideAtlas? Definition and What Research Reports
PeptideAtlas is a publicly accessible database of peptides that have actually been observed in tandem mass spectrometry experiments, built by reprocessing raw proteomics data through a uniform pipeline and mapping the peptide identifications back to an organism's genome. It is a data resource, not a drug, supplement, or research compound, so no dose or route is associated with the term. Published builds cover human plasma, Candida albicans, maize, pig, horse, and Arabidopsis, and researchers have used such peptide evidence to judge which predicted proteins are genuinely observed.
Definition
PeptideAtlas is a public, multi-organism database of peptides that have actually been observed in tandem mass spectrometry (MS/MS) proteomics experiments. Rather than listing proteins that a genome is predicted to encode, it catalogues the peptide sequences that instruments have detected in real samples, after raw data from many contributing laboratories are reprocessed through a single, uniform analysis and statistical-validation pipeline and the surviving peptide identifications are mapped back to the genome of the organism studied; the project and this design were described by researchers in a 2006 report in Nucleic Acids Research (PMID 16381952). In everyday proteomics usage, saying that a peptide or protein is "in PeptideAtlas" is shorthand for saying that experimental mass-spectrometry evidence for it exists and has passed a common quality threshold.
What class of thing is it — and what it is not
PeptideAtlas is a bioinformatics resource: a curated database plus the analysis pipeline and web interface that serve it. It is not a molecule, a peptide product, a compound, or anything that is administered to humans or animals. There is therefore no dose, route, schedule, or safety profile attached to the term itself. The "peptides" in PeptideAtlas are tryptic and other proteolytic fragments generated in the laboratory when proteins from a sample are digested before mass-spectrometry analysis — they are analytical fragments used as evidence of the parent protein, not synthesised research peptides.
This page is for educational purposes only and is not medical advice; consult a licensed physician or qualified clinician for any question about health, diagnosis, or treatment.
Where the data come from
The contents of PeptideAtlas originate as raw instrument files donated by proteomics laboratories worldwide. Those files are not simply archived: the project's stated approach was to reprocess submitted data through a consistent search and validation workflow so that identifications from different labs, instruments, and search engines could be compared on the same footing, with the results assembled into genome-mapped "builds" (PMID 16381952). Builds are organised by organism and often by sample type, and are periodically rebuilt as new datasets accumulate.
Typical uses of the term in research writing
- Evidence of observation. Authors cite PeptideAtlas to show that a protein has been detected at least once by MS, as opposed to being predicted only from gene models.
- Assay design. Because the database records which peptides are repeatedly observed, it is commonly consulted when selecting proteotypic peptides for targeted quantitative methods.
- The "dark proteome". Researchers use the absence of a protein from a build to define the set of predicted proteins never yet observed — a framing used explicitly in a 2024 analysis of the Arabidopsis proteome, its post-translational modifications, and its unobserved (dark) fraction in PeptideAtlas (PMID 38104260).
- Annotation decisions. Peptide-level evidence is used as an input when deciding whether a candidate reading frame should be annotated as a real protein.
What the published literature reports
One of the earliest and most cited builds was the Human Plasma PeptideAtlas, described in Proteomics in 2005, in which researchers combined multiple plasma proteomics datasets into a single catalogue of peptides observed in human blood plasma (PMID 16052627). Plasma is analytically difficult because a small number of abundant proteins dominate the signal, and the study framed a pooled, uniformly processed peptide catalogue as a way to summarise what had collectively been detected across many separate experiments.
Subsequent publications extended the same approach to other organisms. A Candida albicans PeptideAtlas was reported in the Journal of Proteomics in 2014 as a community resource for that fungal pathogen (PMID 23811049). An Equine PeptideAtlas was published in Proteomics in 2014 and presented as a resource for developing proteomics-based veterinary research (PMID 24436130), and a Pig PeptideAtlas followed in Proteomics in 2016, described by its authors as a resource for systems biology in animal production and biomedicine (PMID 26699206). On the plant side, the Zea mays PeptideAtlas was reported in the Journal of Proteome Research in 2024 as a new maize community resource (PMID 39101213), complementing the Arabidopsis build, whose 2024 report also catalogued detected post-translational modifications alongside the observed proteome (PMID 38104260).
Published builds referenced on this page
| Build / organism | Journal and year | How the report framed it |
|---|---|---|
| Human plasma | Proteomics, 2005 (PMID 16052627) | A pooled catalogue of peptides observed across multiple human plasma proteomics datasets (PMID 16052627). |
| Candida albicans | Journal of Proteomics, 2014 (PMID 23811049) | A PeptideAtlas build for a fungal pathogen (PMID 23811049). |
| Horse (equine) | Proteomics, 2014 (PMID 24436130) | Presented as a resource for proteomics-based veterinary research (PMID 24436130). |
| Pig | Proteomics, 2016 (PMID 26699206) | Described as a resource for systems biology in animal production and biomedicine (PMID 26699206). |
| Maize (Zea mays) | Journal of Proteome Research, 2024 (PMID 39101213) | Introduced as a new maize community resource (PMID 39101213). |
| Arabidopsis | Journal of Proteome Research, 2024 (PMID 38104260) | Reported detection of the proteome and its post-translational modifications, and characterised the unobserved "dark" proteome (PMID 38104260). |
Doing the math on a vial? The PeptideU app does reconstitution, units and dilution for you.
Try it freePeptide evidence and the edges of the proteome
A recurring theme in the recent literature is how strict the evidence should be before a predicted product is called a protein. A 2025 preprint addressed exactly this question, with researchers reporting on the use of high-quality peptide evidence for annotating non-canonical open reading frames as human proteins (PMID 39314370). The same broader question — which small translated products belong in the human protein catalogue — was taken up in a 2026 Nature report on expanding the human proteome with microproteins and peptideins (PMID 42092140). Both illustrate why uniformly processed peptide catalogues matter: the decision to add an entry to a reference proteome depends on how confidently the underlying spectra were matched.
Limitations noted in the literature
Absence from a build does not mean a protein does not exist. Mass spectrometry is biased toward abundant, easily digested, and readily ionised proteins, which is why the Arabidopsis analysis treated the never-observed fraction as its own object of study rather than as evidence of non-existence (PMID 38104260). Coverage also depends on what laboratories have contributed: the human plasma build reflected the plasma datasets available to it at the time (PMID 16052627), and organism-specific builds reflect the sampling depth of their respective communities (PMID 39101213).
Related terms
- Proteome — the full set of proteins expressed by an organism, tissue, or cell state.
- Proteotypic peptide — a peptide that reliably and uniquely identifies its parent protein in MS experiments.
- Tandem mass spectrometry (MS/MS) — the fragmentation-based technique that generates the spectra underlying peptide identifications.
- Dark proteome — predicted proteins for which no observed peptide evidence has yet been recorded (PMID 38104260).
- Non-canonical open reading frame — a reading frame outside standard gene annotation whose protein status depends on peptide-level evidence (PMID 39314370).
Tracking research? Log entries with dates, lots and notes — records, never plans.
Get the appReferences
- The PeptideAtlas project (Nucleic Acids Research, 2006)
- Human Plasma PeptideAtlas (Proteomics, 2005)
- A Candida albicans PeptideAtlas (Journal of Proteomics, 2014)
- The Equine PeptideAtlas: a resource for developing proteomics-based veterinary research (Proteomics, 2014)
- The Pig PeptideAtlas: A resource for systems biology in animal production and biomedicine (Proteomics, 2016)
- The Zea mays PeptideAtlas: A New Maize Community Resource (Journal of Proteome Research, 2024)
- Detection of the Arabidopsis Proteome and Its Post-translational Modifications and the Nature of the Unobserved (Dark) Proteome in PeptideAtlas (Journal of Proteome Research, 2024)
- High-quality peptide evidence for annotating non-canonical open reading frames as human proteins (bioRxiv, 2025)
- Expanding the human proteome with microproteins and peptideins (Nature, 2026)
Frequently asked questions
Is PeptideAtlas a peptide?▾
No. PeptideAtlas is a database, not a molecule. It catalogues peptide sequences that were observed in tandem mass spectrometry experiments after raw data were reprocessed through a uniform pipeline and mapped to the genome, as described in the project report (PMID 16381952). Nothing in it is administered to people or animals, so no dose, route, or schedule is associated with the term.right
Where does the data in PeptideAtlas come from?▾
It comes from raw mass-spectrometry files contributed by proteomics laboratories. Rather than archiving them as-is, the project reprocessed submitted datasets through a consistent search and statistical-validation workflow so identifications from different labs and instruments could be compared, then assembled genome-mapped builds (PMID 16381952). Builds are organised by organism and are periodically rebuilt as new datasets accumulate.
Which organisms have published PeptideAtlas builds?▾
Published builds span humans and many model and agricultural species. Reports include the Human Plasma PeptideAtlas (PMID 16052627), a Candida albicans build (PMID 23811049), an Equine build framed as a veterinary research resource (PMID 24436130), a Pig build for systems biology in animal production and biomedicine (PMID 26699206), and a Zea mays maize community resource (PMID 39101213).
What is the "dark proteome" in this context?▾
The dark proteome refers to proteins a genome is predicted to encode but for which no observed peptide evidence has been recorded. A 2024 Journal of Proteome Research analysis reported on the detected Arabidopsis proteome and its post-translational modifications and explicitly characterised the unobserved, or dark, fraction in PeptideAtlas (PMID 38104260). Absence of evidence reflects detection limits, not proof of non-existence.
Why is peptide evidence used to decide what counts as a protein?▾
Because gene models can predict reading frames that may never be translated. Researchers reported on using high-quality peptide evidence to decide which non-canonical open reading frames should be annotated as human proteins (PMID 39314370), and a 2026 Nature report addressed expanding the human proteome with microproteins and peptideins (PMID 42092140). Confidence in the underlying spectral matches drives annotation decisions.
What is the Human Plasma PeptideAtlas?▾
It was an early organism- and sample-specific build described in Proteomics in 2005, in which multiple plasma proteomics datasets were combined into a single catalogue of peptides observed in human blood plasma (PMID 16052627). Plasma is analytically challenging because a few abundant proteins dominate the signal, so pooling uniformly processed data summarised what had collectively been detected.
Does inclusion in PeptideAtlas say anything about a compound's safety or effects?▾
No. The resource records analytical observations of protein-derived peptide fragments in laboratory samples; it does not evaluate biological activity, safety, or clinical outcomes. The published build reports describe data coverage and community use, for example in veterinary research (PMID 24436130) and animal production and biomedicine (PMID 26699206). This answer is educational only and is not medical advice.
Track it. Calculate it. Actually understand it.
References
This page summarises published research for education — it is not medical advice, and nothing here is a recommendation to use, purchase, or dose any substance. Study parameters described are what researchers reported, not instructions. Consult a qualified clinician before any health decision.