A peptide's name alone does not guarantee a single, well-defined substance. Many entries sold under one name are fragments, salts, metal complexes, mixtures, or poorly standardized material. The atlas resolves identity through authoritative databases rather than commercial claims.
The full method is documented in Identity and structure assets. To see how identity connects to evidence, read How to read peptide evidence.
Sequence sources
The atlas draws sequences from five types of authoritative sources: DrugBank and FDA/EMA labels for approved drugs, UniProt for proteins, PubChem Compound summaries, and ClinicalTrials.gov for investigational candidates.
For semaglutide, the monograph records a 31-amino-acid sequence with two key modifications: an Aib substitution at position 2 for DPP-IV resistance, and a C18 fatty diacid linked via a hydrophilic spacer to Lys26 for albumin binding. The sequence is verified against PubChem CID 56843331.
For BPC-157, the sequence is Gly-Glu-Pro-Pro-Pro-Gly-Lys-Pro-Ala-Asp-Asp-Ala-Gly-Leu-Val — a linear 15-residue peptide verified through PubChem CID 9941957.
PubChem structures and their limits
For small peptides (<40 residues), PubChem typically records a defined chemical structure. The atlas retrieves a 2D PNG through PubChem's image service, accompanied by a provenance file recording the stable ID, retrieval URL, date, and SHA-256 checksum.
Two-dimensional depictions have important limitations. They flatten three-dimensional structure entirely. Disulfide bridges — critical for oxytocin, ziconotide, and linaclotide — are drawn as S-S bonds without showing spatial arrangement. Cyclic peptides such as cyclosporine and daptomycin appear as planar cycles rather than complex macrocyclic folds. As the identity policy states: a depiction is not an experimentally determined conformation.
For peptides larger than 40 residues, or those without a stable defined small-molecule structure — dulaglutide, mecasermin, follistatin-344 — no 2D depiction is retrieved. These entries are marked not_applicable in the structure manifest.
Salts, conjugates, and mixtures
Many catalog entries are not single, neutral peptides. The identity policy records the specific form:
Salts: peptide·HCl or acetate salts such as glatiramer acetate. The PubChem record for the free base is used, and the salt form is noted.
Fatty-acid conjugates: semaglutide's C18 diacid, liraglutide's C16 palmitoyl, and palmitoyl cosmetic peptides. The conjugate is integral to the molecule.
Metal complexes: GHK-Cu forms a copper(II) complex. PubChem CID 133697840 represents a bis(GHK)-copper (2:1) species — not the 1:1 complex often intended by the name.
Mixtures: thymalin and cerebrolysin are undefined peptide mixtures from tissue extracts. No single CID or sequence applies.
Random copolymers: glatiramer acetate has no defined sequence.
Stereochemistry and modifications
PubChem 2D depictions can encode assigned stereochemistry through wedge and hash bonds, but they remain flattened representations. Many catalog entries contain D-amino acids that are critical for activity — octreotide, ipamorelin, and cetrorelix all employ D-amino acid substitutions for proteolytic stability. The identity table records these in the modifications_conjugates field. For example, setmelanotide is an 8-residue cyclic peptide with D-Ala at position 3 and D-Phe at position 5, plus N-terminal acetylation and a C-terminal cysteinamide. A 2D depiction alone cannot guarantee that every relevant stereocentre, atropisomer, or product-specific configuration has been resolved.
Boundary cases with ambiguous identity
Several entries have verified PubChem CIDs that do not correspond to the exact entity expected. The identity policy records these explicitly:
TB-500 is marketed as a thymosin beta-4 fragment, but identity is often ambiguous between full-length Tβ4 and a heptapeptide fragment. The TB-500 monograph records this caveat in its identity table.
Modified GRF (1-29) and CJC-1295 are sequence variants of GHRH whose exact marketed products may differ.
IGF-1 LR3 and Des(1-3) IGF-1 are modified protein sequences not captured by a single CID.
PEG-MGF is a PEGylated protein with no defined small-molecule structure.
The GHK-Cu monograph illustrates a particularly subtle case: PubChem CID 133697840 represents a bis(GHK)-copper (2:1) species, not the 1:1 complex often intended by the name. The CID is retained as a registry pointer, but its depiction is not used as the generic entry's asset. The caveat field explains this directly.
Asset management and provenance
Each retrieved structure image in assets/structures/ has a corresponding provenance file recording the database, stable ID, retrieval URL, date, SHA-256 checksum, and license terms. The manifest CSV (assets/structures/manifest.csv) aggregates these fields. Checksums are verified before commit. The data in data/identities.csv records the same verification dates — all verified as of 2026-08-06 across all catalog entries.
Refresh policy
Identities and structure assets are verified as of 2026-08-06. PubChem records may be updated; sequences from DrugBank and UniProt should be rechecked for major revisions. Mixture-based entries — thymalin, cerebrolysin, glatiramer — have no CID to refresh, so their identity records are inherently less machine-verifiable than entries with stable small-molecule references.
Why identity verification matters for evidence
Identity verification is not a side issue — it directly affects whether a study's results can be attributed to the substance named. A clinical trial tests a specific, well-characterized product. Material sold under the same name in the research market may have different purity, composition, stereochemistry, or modifications. Even a study of an investigational product does not automatically validate material sold elsewhere.
The identity methodology page provides the complete asset management rules, database list, and unresolved boundary examples for every catalog entry.