What Is Amino Acid Sequence? Definition and What Research Reports
An amino acid sequence is the order in which amino acids are linked together in a peptide or protein chain, read from the N-terminus to the C-terminus. It is also called primary structure. The sequence is what distinguishes one peptide from another, and published work has reported that changing the order of the same amino acids can change how a molecule assembles. This glossary entry defines the term, explains how it is used and misused in peptide research, and summarises cited literature.
Plain-language definition
An amino acid sequence is simply the order of the amino acids strung together in a peptide or protein. Amino acids are small molecules that link end-to-end like beads on a string; the sequence is the list of which bead comes first, second, third and so on. Two peptides can contain exactly the same amino acids in different orders and still be different molecules with different behaviour. When a paper, a certificate of analysis or a database entry identifies a peptide, the sequence is normally the first identifying feature given, because it is the most basic description of what the molecule is.
The term in biochemical language
In biochemistry, the amino acid sequence is also called the primary structure of a polypeptide. Amino acid residues are joined by peptide bonds, which are covalent amide links formed between the carboxyl group of one residue and the amino group of the next. Because that bond has a direction, the chain has two distinguishable ends: the N-terminus (free amino group) and the C-terminus (free carboxyl group). By universal convention, sequences are written and read from the N-terminus to the C-terminus, so writing a sequence backwards describes a different molecule.
How sequences are written
Two notations are standard, and both describe the same chain:
| Notation | Example | Typical use |
|---|---|---|
| Three-letter code | Gly-His-Lys | Chemistry papers, synthesis schemes, figures |
| One-letter code | GHK | Databases, alignments, bioinformatics files |
Primary structure is distinguished from higher levels of organisation: secondary structure (local folds such as helices and sheets), tertiary structure (the three-dimensional shape of one chain) and quaternary structure (how several chains associate). Sequence is the input; the higher levels are, in large part, consequences of it, together with the surrounding environment.
Why sequence order matters: what studies report
The idea that order — not merely composition — carries information has been examined experimentally. A 2019 study in Chemical Communications investigated the role of amino acid sequence order in the assembly and function of the amyloid-β core region, and the researchers reported that the arrangement of residues, rather than their identity alone, shaped how the segments assembled (PMID 31276123). In other words, rearranging the same set of residues was reported to alter assembly behaviour. That finding is a useful anchor for a glossary definition: it illustrates why the term "sequence" is defined as an ordered list and not as a set of ingredients.
Doing the math on a vial? The PeptideU app does reconstitution, units and dilution for you.
Try it freeSequence comparison, alignment and annotation
Much of what is known about any given peptide or protein comes from comparing its sequence with others. Sequence alignment is the process of lining up two or more sequences so that corresponding positions sit in the same column, inserting gaps where residues appear to have been added or lost over evolutionary time. Alignments underpin homology searching, conservation analysis and structure prediction.
A 2015 data article in Data in Brief presented an amino acid sequence alignment of vertebrate CAPN3/calpain-3/p94 across species, making the aligned sequences available as a reference dataset (PMID 26958593). The study is an example of a common use of the term in practice: sequences from multiple organisms were compiled and aligned so that conserved and divergent positions could be inspected directly.
Alignment quality is itself an active research question. A 2022 paper in Proteins examined whether predicted protein structure could be used to improve alignment, and the researchers reported that incorporating structure prediction improved the quality of amino-acid sequence alignment (PMID 35754316). The study is relevant to the definition because it shows that sequence, while foundational, is often interpreted alongside structural information rather than in isolation.
How the term is used in peptide research
In peptide literature, "amino acid sequence" typically appears in several recurring contexts:
- Identity. The sequence, written N-to-C, is the primary identifier of a synthetic peptide, usually accompanied by molecular formula and monoisotopic mass.
- Design. Analogues are described as substitutions, deletions or insertions relative to a parent sequence — for example, replacing one residue at a defined position.
- Fragments. Many studied peptides are defined as numbered fragments of a larger protein, where the numbers refer to positions in the parent sequence.
- Conservation. Cross-species alignments, such as the vertebrate CAPN3 dataset described above, are used to identify positions that have remained unchanged (PMID 26958593).
- Structure–activity discussion. Papers relate sequence changes to changes in folding, aggregation or binding, as in the amyloid-β core work (PMID 31276123).
Tracking research? Log entries with dates, lots and notes — records, never plans.
Get the appWhere the term is misused
Several misuses appear frequently outside the peer-reviewed literature:
- Treating sequence as a complete specification. A sequence does not describe terminal modifications (such as N-terminal acetylation or C-terminal amidation), stereochemistry (L- versus D-amino acids), cyclisation, disulfide pairing, glycosylation, salt form, counter-ion content or purity. Two materials can share a written sequence and still differ chemically.
- Equating sequence similarity with equivalent behaviour. Alignment shows positional correspondence, not functional identity. The 2022 alignment study itself treated alignment as a technical problem with variable quality rather than a definitive statement about function (PMID 35754316).
- Reversing or scrambling the order casually. Because order carries information, a retro sequence or a scrambled sequence is a different molecule; scrambled sequences are in fact often used as controls precisely because order matters (PMID 31276123).
- Using "sequence verified" as a proxy for quality. Confirming a sequence by mass spectrometry or sequencing addresses identity, not sterility, endotoxin content or impurity profile.
Regulatory and labelling context
Sequence length also carries regulatory meaning. Under United States law, the Food and Drug Administration has defined a peptide as a polymer composed of 40 or fewer amino acids, with larger polymers treated as proteins and regulated as biological products; that boundary determines which approval pathway a molecule follows. Separately, many synthetic peptides are distributed labelled "research use only," which is a statement about permitted distribution and labelling rather than an indication of safety or efficacy. Nothing on this page is legal advice.
Want the full course? Every compound, evidence-graded and cited, inside PeptideU.
Start learning freeRelated terms
| Term | Relationship to amino acid sequence |
|---|---|
| Primary structure | Synonym for the ordered list of residues in a chain |
| Residue | A single amino acid unit once incorporated into the chain |
| N-terminus / C-terminus | The two ends that give the sequence its reading direction |
| Peptide bond | The covalent amide link joining consecutive residues |
| Sequence alignment | Method of lining up two or more sequences position by position |
| Homology / conservation | Inferences drawn from similarity between aligned sequences |
| Analogue | A molecule differing from a parent sequence by defined changes |
| Motif | A short recurring sequence pattern associated with a property |
What sequence alone does not tell us
Even a perfectly known sequence leaves open questions about conformation, stability, aggregation state and interaction partners. The 2022 Proteins study illustrates the point from the computational side: the researchers reported that adding predicted structural information improved alignment quality, implying that sequence data alone were not the optimal basis for the comparison (PMID 35754316). On the experimental side, the amyloid-β core work reported that assembly and function depended on the ordering of residues, which is information that a simple amino acid composition table would not capture (PMID 31276123).
This page is for educational purposes only and is not medical advice; consult a licensed physician for any question about health, medicines or research materials. It describes how a term is defined and used in published literature and does not describe or endorse any use of any compound in humans.
Doing the math on a vial? The PeptideU app does reconstitution, units and dilution for you.
Try it freeReferences
- Protein structure prediction improves the quality of amino-acid sequence alignment (Proteins, 2022)
- Unravelling the role of amino acid sequence order in the assembly and function of the amyloid-β core (Chemical Communications, 2019)
- Amino acid sequence alignment of vertebrate CAPN3/calpain-3/p94 (Data in Brief, 2015)
Frequently asked questions
What is an amino acid sequence in simple terms?▾
It is the order of amino acids linked together in a peptide or protein chain, read from the N-terminus to the C-terminus. It is also called primary structure. The sequence is the most basic identifier of a molecule, and published work reported that the order of residues, not only which residues are present, influenced how segments assembled (PMID 31276123).
Does the order of amino acids actually matter?▾
Yes. A 2019 study examined the role of amino acid sequence order in the assembly and function of the amyloid-β core, and the researchers reported that the arrangement of residues shaped assembly behaviour (PMID 31276123). That is why scrambled-sequence peptides are commonly used as experimental controls: they share composition but not order.
What is sequence alignment?▾
Alignment lines up two or more sequences so corresponding positions sit in the same column, with gaps inserted where residues appear to have been added or lost. A 2015 data article published an alignment of vertebrate CAPN3/calpain-3/p94 sequences as a reference dataset, allowing conserved and divergent positions to be inspected directly (PMID 26958593).
Is sequence enough to predict structure?▾
Not on its own. A 2022 paper in Proteins examined whether predicted protein structure could improve alignment, and the study reported that incorporating structure prediction improved the quality of amino-acid sequence alignment (PMID 35754316). That result implies sequence data alone were not the strongest basis for comparison, and structural information added value.
Do two peptides with the same sequence have to be identical?▾
No. A written sequence does not capture terminal modifications such as acetylation or amidation, stereochemistry, cyclisation, disulfide pairing, salt form or purity. Materials sharing a sequence string can differ chemically. Sequence confirmation addresses identity only; it says nothing about impurity profile, sterility or endotoxin content.
How are amino acid sequences written down?▾
Sequences are written from the N-terminus to the C-terminus using either three-letter codes (Gly-His-Lys) or one-letter codes (GHK). Databases and alignment files generally use the one-letter form, as in the vertebrate CAPN3 alignment dataset (PMID 26958593), while synthesis schemes and chemistry figures often use three-letter notation.
What does sequence length have to do with regulation?▾
In the United States, the FDA has defined a peptide as a polymer of 40 or fewer amino acids, with longer polymers treated as proteins and regulated as biological products, which affects the approval pathway. Separately, "research use only" labelling describes permitted distribution, not safety or efficacy. This is not legal advice.
Track it. Calculate it. Actually understand it.
References
This page summarises published research for education — it is not medical advice, and nothing here is a recommendation to use, purchase, or dose any substance. Study parameters described are what researchers reported, not instructions. Consult a qualified clinician before any health decision.