→
→
→
Amino Acids, Peptides, and Proteins
Amino Acids, Peptides, and Proteins
Amino acids are the building blocks that combine to form peptides and proteins.
Amino acids are the building blocks that combine to form peptides and proteins. Every amino acid is defined by two functional groups — an amino group and a carboxyl group — attached to the same central carbon, along with a hydrogen atom and a variable side chain that gives each amino acid its identity. This article covers amino acid structure, their behavior in solution, their stereochemistry, how they're classified, and how they link together to form proteins.
Key Takeaways
An alpha-amino acid has an amino group, a carboxyl group, a hydrogen atom, and a variable R group (side chain) all bonded to the alpha carbon.
Amino acids are amphoteric and predominantly exist as zwitterions (protonated amino group, deprotonated carboxylate) at physiological pH (~7).
The alpha carbon is chiral in nearly all amino acids; glycine (R group = H) is the sole achiral exception.
Eukaryotic amino acids are almost all L-configuration (Fischer projection, amino group on the left, relative to glyceraldehyde). L/D is distinct from R/S (CIP absolute configuration); most L-amino acids are S, but cysteine is L yet R, due to sulfur's CIP priority.
There are 21 proteogenic amino acids in eukaryotes: the standard 20 plus selenocysteine, incorporated during translation via UGA stop-codon recoding directed by a SECIS element.
Amino acids fall into five side-chain categories: nonpolar/nonaromatic, aromatic, polar uncharged, negatively charged (acidic), and positively charged (basic) — each shaping hydrophobicity, hydrogen bonding, and charge at physiological pH.
Peptide bonds form via condensation (water lost) between one amino acid's amino group and another's carboxyl group, building dipeptides and, ultimately, polypeptides.
Hydrolysis (addition of water) reverses peptide bond formation and drives protein digestion.
Alpha-Amino Acid Structure
An alpha-amino acid is defined by four groups bonded to a single central carbon, called the alpha carbon:
An amino group
A carboxyl group
A hydrogen atom
A variable side chain, represented by "R"
The R group is what distinguishes one amino acid from another — everything else in this core structure is shared across all amino acids.
Amphoteric Character and the Zwitterion
Because an amino acid contains both an acidic group (the carboxyl group) and a basic group (the amino group), amino acids are amphoteric — they can act as either an acid or a base depending on their environment.
In a neutral solution, around pH 7, amino acids predominantly exist as a zwitterion: a molecule that carries both a positive and a negative charge while remaining overall electrically neutral. In this form, the amino group gains a proton and becomes positively charged, while the carboxyl group loses a proton and becomes a negatively charged carboxylate group. This balance of charges is what stabilizes amino acids in aqueous environments at physiological pH.
Chirality of the Alpha Carbon
The alpha carbon is chiral in almost every amino acid — it's bonded to four different groups, making the molecule non-superimposable on its mirror image.
Glycine is the one exception: its side chain is simply a hydrogen atom, so the alpha carbon ends up bonded to two identical hydrogen atoms rather than four different groups. Glycine is therefore achiral and does not exhibit optical activity.
Because the rest of the amino acids are chiral, they can rotate plane-polarized light, and their configuration can be described in two different ways: L or D, and R or S.
L and D Configuration
In eukaryotes, almost all amino acids occur in the L-configuration. The L/D labeling system is based on a Fischer projection of the amino acid compared against a historical reference molecule, glyceraldehyde: if the amino group is drawn on the left side of the projection, the amino acid is classified as L; if it's drawn on the right, it's classified as D.
L/D vs. R/S Configuration
L and D are not the same thing as R and S. R and S describe the absolute configuration at the alpha carbon, determined using the Cahn–Ingold–Prelog (CIP) priority rules — an entirely different system from the Fischer-projection-based L/D labels.
When the R/S rules are applied to amino acids, almost all naturally occurring L-amino acids turn out to have an S configuration at the alpha carbon.
MCAT Callout — Cysteine's R/S Exception: Cysteine is classified as an L-amino acid, but it has an R absolute configuration — the opposite of the usual L-amino acid pattern. This happens because cysteine's side chain contains a sulfur atom, and sulfur has a higher atomic number than oxygen. That higher atomic number shifts sulfur's priority rank under the CIP rules, which changes the overall priority order at the alpha carbon and flips the assigned configuration from S to R — even though cysteine's L/D classification (based on the separate Fischer-projection convention) is unaffected and stays L.
In summary, in eukaryotic proteins: amino acids are almost always L-amino acids, most have an S configuration at the alpha carbon, and cysteine is the exception — L, but R.
The 21 Eukaryotic Proteogenic Amino Acids
There are 21 proteogenic amino acids in eukaryotes — the amino acids that are genetically encoded. That's the standard 20 amino acids plus selenocysteine, considered the 21st proteogenic amino acid. Selenocysteine is directly incorporated into proteins during translation through a special mechanism.
MCAT Callout — How Selenocysteine Is Incorporated: Selenocysteine is inserted by recoding an in-frame UGA stop codon — normally a "stop" signal — into a sense codon for selenocysteine. This recoding is directed by a specific mRNA structural element called a SECIS (selenocysteine insertion sequence) element.
Classifying Amino Acids by Side Chain
The 21 proteogenic amino acids fall into five categories based on the chemical character of their side chains:
Nonpolar, nonaromatic — alanine, valine, leucine, isoleucine, methionine, glycine, and proline. Their side chains are primarily hydrocarbon-based, making them generally hydrophobic; they're often buried in the interior of a protein to avoid contact with water. Proline is a special case: its side chain loops back and bonds to the amino group, giving it a rigid, cyclic structure.
Aromatic — phenylalanine, tyrosine, and tryptophan. These side chains contain aromatic rings. Phenylalanine and tryptophan are largely hydrophobic, while tyrosine carries a polar hydroxyl group on its aromatic ring that makes it somewhat more hydrophilic.
Polar, uncharged — serine, threonine, asparagine, glutamine, and cysteine. Their side chains contain oxygen, nitrogen, or sulfur atoms, which allow hydrogen bonding, but they carry no net charge at physiological pH.
Negatively charged (acidic) — aspartic acid and glutamic acid. At physiological pH, these side chains have lost a proton from their carboxyl groups, leaving a negative charge. They're highly hydrophilic and often interact with positively charged residues or ions.
Positively charged (basic) — lysine, arginine, and histidine. These side chains contain protonated amino groups, giving them a positive charge under physiological conditions. They're hydrophilic and often participate in binding negatively charged molecules, such as DNA or phosphate groups.
MCAT Callout — Histidine's Borderline Charge: Histidine's imidazole side chain has a pKa around 6, close to physiological pH (~7.4). Unlike lysine and arginine, which stay fully protonated at physiological pH, histidine's side chain is often mostly unprotonated at that pH — it's grouped as "basic" by convention and by its proton-shuttling role, but its actual charge state is more variable than the other two basic amino acids.
MCAT Callout — Why Side Chain Chemistry Matters: The side chain's chemical nature shapes how an amino acid behaves in water, how it interacts with other molecules, and how a protein folds into its three-dimensional structure.
Peptide Bond Formation
Amino acids link together through condensation reactions — reactions that join two molecules while releasing a small molecule, in this case water — to form peptide bonds.
When two amino acids react, the amino group of one attacks the carbonyl carbon of the other's carboxyl group. As this new bond forms, a molecule of water is eliminated, producing an amide bond, also called a peptide bond.
One amino acid contributes its free amino group, called the amino terminus (or N-terminus).
The other contributes its free carboxyl group, called the carboxy terminus (or C-terminus).
The nitrogen from the amino group attacks the carbonyl carbon, forming the new bond and releasing water.
The resulting two-amino-acid molecule is a dipeptide — two amino acids joined by one peptide bond.
This same linking pattern continues to build longer chains of amino acids, called polypeptides, which are the primary structural units of proteins.
Hydrolysis: Breaking Peptide Bonds
Peptide bond formation is reversible. Hydrolysis — the addition of water — breaks peptide bonds back down, and it can occur under either acidic or basic conditions. Hydrolysis is the reverse of condensation, and it plays a critical biological role, particularly in the breakdown of proteins during digestion.
Together, the ability of amino acids to form peptide bonds through condensation and to be broken back apart through hydrolysis is the chemical basis for the dynamic structure and function of proteins in living organisms.
Common MCAT Mistakes
Treating L/D and R/S as the same labeling system. They're independent systems — L/D comes from a Fischer projection comparison to glyceraldehyde, while R/S comes from CIP priority rules at the alpha carbon. Most L-amino acids happen to be S, but cysteine is L yet R because sulfur's atomic number changes its CIP priority.
Forgetting glycine is the one achiral amino acid. Every other amino acid has four different groups on its alpha carbon and is chiral; glycine's side chain is just a hydrogen, giving it two identical substituents and no chirality.
Undercounting the proteogenic amino acids as 20 instead of 21. Selenocysteine is the 21st, inserted co-translationally by recoding a UGA stop codon via a SECIS element — it's easy to forget because it isn't one of the "standard 20."
Assuming histidine is fully protonated like lysine and arginine. Histidine's imidazole side chain has a pKa near 6, close to physiological pH, so it's often mostly unprotonated in vivo — its "basic" classification reflects its chemistry and proton-shuttling role, not a guarantee of full positive charge at pH 7.4.
MCAT-Style Concept Check
Question: Cysteine is classified as an L-amino acid but has an R absolute configuration, unlike most other L-amino acids, which are S. What best explains this?
A) Cysteine's amino group is drawn on the right side of its Fischer projection, unlike other L-amino acids
B) Sulfur's higher atomic number shifts its CIP priority rank, changing the overall priority order at the alpha carbon and flipping S to R
C) Cysteine is achiral, so R and S labels are assigned arbitrarily
D) The L/D and R/S systems disagree for cysteine because cysteine does not form a zwitterion
Answer: B
Explanation: Cysteine's side chain contains a sulfur atom, and sulfur outranks oxygen in atomic number. Under the CIP rules, that higher atomic number gives the sulfur-containing side chain a higher priority than it would have with an oxygen-containing group, which reorders the four substituent priorities at the alpha carbon and flips the assigned configuration from S to R. (A) is wrong — cysteine's L classification still follows the standard Fischer-projection convention (amino group on the left); L/D isn't affected by CIP priority. (C) is wrong — cysteine's alpha carbon has four different substituents and is chiral, like nearly all other amino acids. (D) is wrong — cysteine does form a zwitterion like other amino acids at physiological pH; that has no bearing on its R/S assignment.
FAQ
What is a zwitterion, and why do amino acids form one at physiological pH?
A zwitterion is a molecule that carries both a positive and a negative charge while remaining overall electrically neutral. Amino acids are amphoteric, so at physiological pH (~7) the amino group gains a proton (becoming positively charged) while the carboxyl group loses one (becoming a negatively charged carboxylate), producing this balanced, stabilized charge state.
Why is glycine the only achiral amino acid?
Chirality requires four different groups on the alpha carbon. Glycine's side chain is just a hydrogen atom, so its alpha carbon ends up bonded to two identical hydrogens rather than four different groups, making it superimposable on its mirror image and therefore achiral.
What are the 21 proteogenic amino acids, and how is the 21st one incorporated?
The 21 proteogenic amino acids are the standard 20 genetically encoded amino acids plus selenocysteine. Selenocysteine is incorporated during translation by recoding an in-frame UGA stop codon into a sense codon, a process directed by an mRNA structural element called a SECIS (selenocysteine insertion sequence) element.
Why is histidine sometimes treated differently from lysine and arginine, even though all three are classified as basic?
Histidine's imidazole side chain has a pKa around 6, close to physiological pH (~7.4), so it's often mostly unprotonated in the body — unlike lysine and arginine, which stay fully protonated at physiological pH. Histidine is still grouped as basic by convention and for its proton-shuttling role, but its actual charge state is more variable.