El contenido de la evidencia se mantiene en inglés.

Overview

This page documents how the atlas resolves the identity, sequence, and chemical structure of each catalog entry. Many entries sold under a single name are not a single well-defined substance; the goal is to record what authoritative databases say, not what any vendor claims.

Three identity caveats

A database record is a registry pointer, not proof of a commercial sample's identity. A two-dimensional depiction cannot establish three-dimensional or formulation context. When authoritative evidence does not resolve a unique entity, the atlas retains the ambiguity instead of inventing a structure.

Identity resolution stack
Six-layer identity resolution stack with an unresolved-identity boundaryEach lower evidence layer narrows what a marketed or common name means. Undefined mixtures and unresolved identities branch to documented boundaries rather than invented structures.marketed / common namepreferred identity + aliasessequence and modificationschemical form / salt / conjugate / complexstable database identifiersasset + provenance sidecareach lower layer narrowswhat the name meansundefined mixture or unresolved identitydocument boundary — do not invent structure
Database identifiers and assets follow identity resolution; they do not silently settle an ambiguous marketed name or sample.
Alternativa textual

Resolution moves from marketed or common name through preferred identity and aliases, sequence and modifications, chemical form, stable identifiers, and asset provenance. Each layer narrows meaning. An undefined mixture or unresolved identity is documented as a boundary, and no structure is invented.

Biological sequence versus chemical structure

Peptides and proteins are linear polymers of amino acids. Their primary structure (sequence) is the order of residues. For small peptides (under 40 residues), PubChem typically records a defined chemical structure: the sequence of atoms and bonds. For larger peptides and proteins, the stored CID corresponds to a simplified atomic representation or a sequence-only record; a 2D structure image cannot capture the three-dimensional fold, disulfide pairing, or post-translational modifications.

The atlas draws the following distinction:

  • Sequence source: an authoritative source for the amino-acid sequence (DrugBank, UniProt, FDA/EMA label, or PubChem Compound summary).

  • Structure source: the from which the atlas retrieves SMILES for a locally generated 2D depiction. The depiction is a 2D rendering of the structural formula, not an experimentally determined conformation.

The fuller identity workflow is also explained in How Peptide Identity Is Verified.

Sequence is not the whole structure
Residue order plus chemical modifiers defines a context-specific chemical entityA sequence card passes through six modifier checks before a context-specific chemical depiction. TB-500, GHK-Cu, thymalin, and cerebrolysin remain labeled as ambiguous where appropriate.SEQUENCEresidue orderMODIFIER CHECKSterminicyclizationcrosslinksstereochemistryconjugatessalts / complexesCHEMICAL ENTITYcontext-specific depictionTB-500 · GHK-Cu · thymalin / cerebrolysinAMBIGUITY RETAINED
Residue order is one identity layer. Chemical form and unresolved mixture boundaries remain visible in both records and assets.
Alternativa textual

Sequence or residue order must be considered with termini, cyclization, crosslinks, stereochemistry, conjugates, and salts or complexes before a context-specific chemical entity can be depicted. TB-500, GHK-Cu, thymalin, and cerebrolysin retain their documented ambiguity.

Authoritative databases consulted

SourceIdentity roleWhat it can establishLimitation
PubChemPrimary structure-asset sourceCID assignment, record title, and the SMILES used for a locally generated depiction, verified through PUG RESTA registry record or depiction does not authenticate a commercial sample
DrugBankDrug identity sourceSequence and pharmacological identity for approved drugsIt does not transfer one product identity to differently marketed material
UniProtProtein identity sourceProtein sequence and family classificationA protein record does not define every modified, conjugated, or marketed form
ChEBIOntology sourceSmall-molecule ontology identifiersAn ontology identifier alone does not establish sample composition
FDA/EMA product labelsProduct-specific official sourceOfficial sequence and composition for approved therapeuticsThe statement remains product- and jurisdiction-specific
ClinicalTrials.gov recordCandidate sequence or identity where the record discloses itPublic records may not disclose enough detail to resolve the entity

For PubChem-derived assets, each retrieved SMILES record and locally drawn 2D SVG has a cached input and checksum provenance sidecar.

Salts, conjugates, complexes, and mixtures

Many catalog entries are not single, neutral peptides:

  • Salts: peptide hydrochloride and acetate salts, including glatiramer acetate. The PubChem record for the free base is used; salt forms are noted.

  • Fatty-acid conjugates: semaglutide (C18 diacid), liraglutide (C16 palmitoyl), and palmitoyl cosmetic peptides. The conjugate is integral to the molecule; the covers the full conjugate.

  • Metal complexes: GHK-Cu, a copper(II) complex. The CID includes the copper ion.

  • Mixtures: thymalin and cerebrolysin, undefined peptide mixtures from tissue extracts. No single CID or sequence applies; no structure image is retrieved.

  • Random copolymers: glatiramer acetate. No defined sequence applies, so no structure image is retrieved.

Stereochemistry

PubChem 2D depictions can encode assigned stereochemistry through wedge/hash bonds and related annotations. They are nevertheless flattened depictions, not three-dimensional conformations, and they do not guarantee that every relevant stereocentre, atropisomer, salt form, or product-specific configuration has been resolved. Many entries contain D-amino acids, including D-Phe, D-Trp, and D-Arg in octreotide, ipamorelin, and cetrorelix, that are critical for activity. Identity review therefore uses the registry record and the documented modifications_conjugates field, not the PNG alone.

Limitations of 2D depictions

  • 2D images flatten three-dimensional structure entirely.

  • Disulfide bridges, including those in oxytocin, ziconotide, and linaclotide, are shown as S—S bonds but the image does not indicate spatial arrangement.

  • Cyclic peptides, including cyclosporine, daptomycin, and bacitracin, are depicted as planar cycles, not as complex macrocyclic folds.

  • Glycosylation in vancomycin and dalbavancin is shown as attached sugars; assigned stereochemistry may be drawn, but a flat image does not establish product composition, conformation, or analytical identity.

  • Retrieval from PubChem uses the PUG REST property endpoint. RDKit draws the returned SMILES locally. This is a depiction, not an analytically determined structure.

  • For proteins and large peptides without a stable defined small-molecule structure, including dulaglutide, mecasermin, and follistatin-344, no 2D depiction is retrieved. These entries are marked not_applicable in assets/structures/manifest.csv.

Registry identity and sample authenticity are different questions. The product quality, authenticity, and testing guide explains why specialized analytical methods and validated specifications are needed to assess material rather than a database record.

Boundary cases explicitly unresolved

EntryAmbiguityAsset decisionEvidence consequence
Modified GRF (1-29) and CJC-1295Sequence variants of GHRH; the exact marketed product may differRetain the variant boundary rather than using one depiction as genericEvidence for one defined variant is not silently transferred to another
TB-500The marketed name may mean full-length thymosin beta-4 or a fragmentKeep the TB-500 identity and asset boundary explicitClaims remain tied to the studied material, not the ambiguous market name
IGF-1 LR3 and Des(1-3) IGF-1Modified protein sequences are not captured by a single CIDDo not assign a single CID-derived depiction as definitiveRegistry evidence cannot resolve the marketed entity by itself
PEG-MGFPEGylated protein without a defined small-molecule structureNo defined small-molecule depictionStructural and evidence claims retain the unresolved conjugate boundary
GHK-CuCID 133697840 represents a bis(GHK)-copper 2:1 species, not the generic 1:1 complex often intended by the nameRetain the CID as a registry pointer but do not use its depiction as the generic entry assetEvidence is not generalized across differently defined complexes
Apraglutide and cagrilintide sequences are not yet fully disclosed in public regulatory filingsRecord the public-identity limit rather than filling the gapClaims cannot rely on an invented sequence or structure

Asset management

Generated 2D and 3D asset classes

ClassFile/provenance requirementRefresh trigger
Generated 2D SVGPubChem SMILES or verified sequence input, source sidecar, checksum, and manifest rowA stable record, sequence, identifier, or depiction input changes
Idealised 3D still or turntableSequence-build provenance plus the required idealised-conformer label and recorded omissionsThe verified sequence, modifications, or recorded approximation changes
Blend stillComponent provenance, shared-scale record, and unresolved-component notesA component identity, scale basis, or omission record changes

RDKit generates the 2D SVG class from a PubChem SMILES record or, where no CID is recorded, from a verified peptide sequence. PyMOL generates the 3D still and turntable class by building an idealised sequence conformer; blend stills place those component models on a shared scale. The 3D class is neither an experimental structure nor a structure prediction. Every displayed 3D asset therefore carries the label Idealised conformer built from sequence; not an experimental or predicted structure. Conjugates, terminal decorations, noncanonical-residue approximations, unresolved blend components, and other omissions are recorded in the adjacent provenance sidecar.

Each retrieved image in assets/structures/ has:

  1. a corresponding <slug>.source.json provenance file recording the database, stable ID, retrieval URL, date, SHA-256, and license terms;

  2. an entry in assets/structures/manifest.csv recording those fields.

Checksums are verified before commit.

Refresh policy

Identities and structure assets are verified as of 2026-08-06. PubChem records may be updated; sequences from DrugBank and UniProt should be rechecked for major revisions. Mixture-based entries—thymalin, cerebrolysin, and glatiramer— have no CID to refresh.

Preguntas

Is a peptide sequence enough to define the marketed substance?

Not always. Termini, cyclization, crosslinks, stereochemistry, conjugates, salts, complexes, and mixture boundaries can distinguish the marketed substance from a residue sequence alone.

Does a PubChem identifier verify a commercial sample?

No. A PubChem identifier points to a registry identity, while sample authenticity and composition require separate product records and suitable analytical evidence.

Why are some entries left unresolved?

They remain unresolved when authoritative evidence does not support one unique sequence or chemical entity. The atlas records that boundary instead of inventing an identity or structure.

How is structure-asset provenance recorded?

Generated asset classes are linked to source sidecars, checksums, and the structure manifest. The refresh policy calls for rechecking when an identity record, sequence, identifier, provenance field, or asset input changes.