Overview
This page documents how the atlas resolves the identity, sequence, and chemical structure of each catalog entry. Many entries sold under a single name are not a single well-defined substance; the goal is to record what authoritative databases say, not what any vendor claims.
Biological sequence versus chemical structure
Peptides and proteins are linear polymers of amino acids. Their primary structure (sequence) is the order of residues. For small peptides (<40 residues), PubChem typically records a defined chemical structure (the sequence of atoms and bonds). For larger peptides and proteins, the stored CID corresponds to a simplified atomic representation or the sequence-only record; a 2D structure image cannot capture the three-dimensional fold, disulfide pairing, or post-translational modifications.
The atlas draws the following distinction:
Sequence source: an authoritative source for the amino-acid sequence (DrugBank, UniProt, FDA/EMA label, or PubChem Compound summary).
Structure source: the PubChem CID from which a 2D depiction was retrieved through PubChem's image service. The depiction is a 2D rendering of the structural formula, not an experimentally determined conformation.
Authoritative databases consulted
PubChem — primary structure-asset source. CID assignment and titles are verified through PUG REST; each 2D PNG is retrieved through PubChem's image service and accompanied by a URL-and-checksum provenance sidecar.
DrugBank — sequence and pharmacological identity for approved drugs.
UniProt — protein sequence and family classification.
ChEBI — small-molecule ontology identifiers.
FDA/EMA product labels — official sequence and composition for approved therapeutics.
ClinicalTrials.gov — sequence/identity for investigational candidates where available.
Salts, conjugates, complexes, and mixtures
Many catalog entries are not single, neutral peptides:
Salts: peptide·HCl, acetate salts (e.g., glatiramer acetate). The PubChem record for the free base is used; salt forms are noted.
Fatty-acid conjugates: semaglutide (C18 diacid), liraglutide (C16 palmitoyl), palmitoyl cosmetic peptides. The conjugate is integral to the molecule; PubChem CID covers the full conjugate.
Metal complexes: GHK-Cu (copper(II) complex). The CID includes the copper ion.
Mixtures: thymalin, cerebrolysin (undefined peptide mixtures from tissue extracts). No single CID or sequence applies; no structure image is retrieved.
Random copolymers: glatiramer acetate. No defined sequence; no structure image is retrieved.
Stereochemistry
PubChem 2D depictions can encode assigned stereochemistry through wedge/hash
bonds and related annotations. They are nevertheless flattened depictions, not
three-dimensional conformations, and they do not guarantee that every relevant
stereocentre, atropisomer, salt form, or product-specific configuration has
been resolved. Many entries contain D-amino acids (e.g., D-Phe, D-Trp, D-Arg
in octreotide, ipamorelin, and cetrorelix) that are critical for activity.
Identity review therefore uses the registry record and the documented
modifications_conjugates field, not the PNG alone.
Limitations of 2D depictions
2D images flatten three-dimensional structure entirely.
Disulfide bridges (e.g., in oxytocin, ziconotide, linaclotide) are shown as S—S bonds but the image does not indicate spatial arrangement.
Cyclic peptides (e.g., cyclosporine, daptomycin, bacitracin) are depicted as planar cycles, not as complex macrocyclic folds.
Glycosylation (vancomycin, dalbavancin) is shown as attached sugars; assigned stereochemistry may be drawn, but a flat image does not establish product composition, conformation, or analytical identity.
Retrieval from PubChem uses the PNG endpoint, which returns a 2D rendering of the PubChem compound record. This is a depiction, not an analytically determined structure.
For proteins and large peptides without a stable defined small-molecule structure (e.g., dulaglutide, mecasermin, follistatin-344), no 2D depiction is retrieved. These entries are marked
not_applicableinassets/structures/manifest.csv.
Boundary cases explicitly unresolved
The following entries have verified CID but the PubChem record may not correspond to the exact entity expected:
Modified GRF (1-29) and CJC-1295: sequence variants of GHRH; exact marketed product may differ.
TB-500: marketed as a thymosin beta-4 fragment, but identity is often ambiguous (full-length Tβ4 vs. fragment).
IGF-1 LR3 and Des(1-3) IGF-1: modified protein sequences not captured by a single CID.
PEG-MGF: PEGylated protein; no defined small-molecule structure.
GHK-Cu: CID 133697840 represents a bis(GHK)-copper 2:1 species, not the generic 1:1 complex often intended by the name; the CID is retained as a registry pointer but its depiction is not used as the generic entry's asset.
Apraglutide and cagrilintide: investigational; sequences not yet fully disclosed in public regulatory filings.
Asset management
Each retrieved image in assets/structures/ has:
A corresponding
<slug>.source.jsonprovenance file (database, stable ID, retrieval URL, date, SHA-256, license terms).An entry in
assets/structures/manifest.csvrecording these fields.
Checksums are verified before commit.
Refresh policy
Identities and structure assets are verified as of 2026-08-06. PubChem records may be updated; sequences from DrugBank and UniProt should be rechecked for major revisions. Mixture-based entries (thymalin, cerebrolysin, glatiramer) have no CID to refresh.