What Is MAOMAO and Why Does FAIR Peptide Toxicity Data Matter?

MAOMAO stands for Metadata-Aware Ontology for Multi-source Annotation Organization, an ontology-guided resource that harmonizes peptide toxicity information drawn from public databases, literature-derived resources and curated datasets (the MAOMAO release paper, 2026). It is not a single database with a single upstream. Its first release integrates 56 sources, spanning public databases, literature-derived resources, supplementary datasets and manually curated repositories (same paper), and consolidates them into 71,857 unique peptide sequences across eight toxicity-related endpoints. The Core Layer holds 574,856 harmonized peptide-endpoint combinations for those sequences. Searches ran on Google Scholar and PubMed between June 2025 and September 2026, so the release reflects resources available through September 2026. Figures of this type vary by source and curation date, and a single-upstream record inside the release carries only the weight of that upstream.
FAIR peptide toxicity data means something narrower than “well-organized data.” It means a toxicity record that carries a persistent identifier, rich provenance metadata and a stated license, so a reader can trace where the claim came from and under what conditions it may be reused.
The analogy: a record is a lab notebook entry, not a verdict
Treat a toxicity annotation the way you would treat a provenance-stamped notebook entry. It records who observed the effect, when, from which source, on which organism, and under which evidence state. It does not certify that a sequence is safe. That distinction shapes every downstream decision in this guide, because a record tells you what was claimed and how well it was documented, not what you have proven about your own candidate.
What FAIR actually requires of a toxicity record
The acronym hides the operative requirements. The original FAIR principles (Wilkinson et al., Scientific Data, 2016) and the GO FAIR overview specify that data carry a persistent identifier (F1), that metadata be rich (F2) and include the identifier of the data they describe (F3), that provenance be detailed (R1.2), that the record follow domain-relevant community standards (R1.3), and that a clear, accessible usage license be stated (R1.1).

MAOMAO’s harmonized fields are the concrete implementation: source identifiers, literature references, organism information, annotation provenance, retrieval dates, file formats and version information where available, with each normalized sequence keyed by a deterministic SHA-256 identifier over the uppercase, whitespace-free sequence. Files ship as CSV, JSON and YAML.
Key Takeaway: A curated toxicity record is a provenance-stamped claim, not a safety verdict.
Why Curated Peptide Toxicity Metadata Changes Candidate Triage

Curated peptide toxicity metadata changes peptide candidate triage because it lets you filter on the strength of the underlying evidence rather than on the presence of a label. A database that stores only “toxic” or “non-toxic” collapses very different observations into one field. MAOMAO separates them. Its evidence states distinguish positive, negative, ambiguous, unlabeled and no-information records, and the negative state splits further: strong negatives were explicitly supported by experimental evidence, while weak negatives originated from randomly selected negative sets or lacked confirmation of the absence of toxicity annotations (the MAOMAO release paper, 2026).
The distribution shows why that distinction carries weight. Across the release, MAOMAO reports 31,153 positive, 91,506 negative, 15,717 ambiguous and 5,044 unlabeled states (MAOMAO’s own documentation of its evidence states, retrieved 2026-07-30). The negative bucket is nearly three times the positive bucket, so a triage rule that treats every negative as clearance is not filtering on evidence. It is filtering on volume.
|
Evidence state |
Count |
What it supports in triage |
|---|---|---|
|
Peptide synteze Positive |
31,153 |
Flag for follow-up; check endpoint and assay context |
|
Negative |
91,506 |
Depends on strong vs weak origin |
|
Ambiguous |
15,717 |
Cannot resolve a go/no-go on its own |
|
Unlabeled |
5,044 |
No usable signal |
Reading one record and deciding what it proves
Take a record carrying an inferred annotation with no assay context attached. The label says negative. The provenance says the call was inferred rather than measured, and no endpoint, concentration or assay system is recorded alongside it. At that point the record does not answer “is this sequence safe to advance.” It answers “has anyone published a sequence-level toxicity observation for this sequence,” and the answer is no.
That is the decision point. A triage rule built on the presence of a label passes the candidate through. A rule built on evidence state routes it to the ambiguous or unlabeled path, where it needs either a literature check or an assay before it earns a place on a shortlist. The second rule costs more per candidate and produces a shortlist you can defend.
The same logic applies when two sources disagree on the endpoint label for one sequence. The curated record does not pick a winner silently. It holds both annotations with their provenance, and the disagreement itself becomes the signal that the sequence needs orthogonal analytical and biological testing rather than a database verdict.
Where the evidence runs out
Endpoint coverage is uneven, and the gap is structural rather than incidental. MAOMAO states that coverage depends on retrievable sequence-level evidence, which produces class imbalance and limited representation of Cytolysis, Embryotoxic and Ichthyotoxic endpoints (the MAOMAO release paper, 2026). The counts make the thinness concrete: Embryotoxic carries 2 positive states and Ichthyotoxic 5, while Cytolysis holds 343 states in total. Compare that with Toxic root at 57,080 states and 13,824 positives, or Hemolytic at 35,960 states and 4,714 positives.
For a triage rule, this means absence of an annotation in a thin endpoint is not evidence of absence. It means the endpoint was rarely recorded. A shortlist filtered on Embryotoxic or Ichthyotoxic data is filtered on almost nothing, and the honest move is to say so before the candidate reaches a program decision.
Licensing is the second boundary. Annotation content and last-update dates are available for 55 of 56 sources, but explicit licensing covers only 15 of 56 (MAOMAO’s own documentation of its evidence states, retrieved 2026-07-30). Reuse terms for the remainder need checking at the source, not assumed from the aggregate.
Key Takeaway: A negative is not one thing. Triage on evidence state and provenance, Winkel not on label presence, and treat thin endpoints as unmeasured rather than clean.
Peptide Toxicity Evidence States and Provenance Fields, Field by Field
A peptide toxicity record is auditable when two things are explicit: what the evidence state is, and where the record came from. MAOMAO structures both. The ontology has a root class, Toxic, with primary classes Cytotoxic, Neurotoxic, Embryotoxic, and Ichthyotoxic, and Cytotoxic carries subclasses Hemolytic, Cytolysis, and Anti mammalian cells (MAOMAO release paper, 2026). Read branch totals with care: the root Toxic overlaps subclass counts, so summing the branches will overstate the root, a limitation the paper’s methodology note states directly.
Identity is handled by hashing the normalized uppercase sequence with SHA-256, which gives the same sequence the same key across layers and sources. Records ship as CSV, JSON, or YAML.
The evidence-state vocabulary in practice
Each state answers a different question for a program. Positive: is there direct or hierarchy-supported toxicity evidence? Negative: is there direct, endpoint-specific non-toxic evidence? Ambiguous: are annotations conflicting or unresolved? Unlabeled: was the sequence retained without an explicit class assignment? No information: is there simply no evidence for this endpoint? (MAOMAO release paper, 2026). Treat these as documentation of what the source said, not as a safety determination.
Provenance fields that make a record auditable
MAOMAO’s harmonized fields let a reviewer check a claim without opening the original source. Source identifier and literature reference satisfy Findable and traceable attribution. Organism and annotation provenance satisfy Reusable context. Retrieval date and version satisfy Accessible currency. File format supports Interoperable exchange. Oer
Licensing is the honest gap. The paper states no unified dataset license in its main text, reports source-level licensing coverage of 15 of 56 sources, and places licensing information and versioning records in the Documentation Layer (MAOMAO release paper, 2026). Check source-level terms before reuse in a regulated workflow.
How FAIR Peptide Toxicity Data Plugs Into Assay Design and Reporting

Curated peptide toxicity metadata tells you which orthogonal confirmation to run next; it does not tell you a candidate is safe. That boundary is the whole point of the layer. A database record summarizes what someone else observed under stated conditions, while demonstrated safety rests on orthogonal analytical and biological testing performed on your material: HPLC purity, MS identity, counterion state, endotoxin by LAL, sterility, and the biological assays your program requires. Each test answers a narrow question. LAL measures endotoxin burden only, not sterility or chemical purity, and counterion mass is not trivial, since the TFA salt can represent roughly 5% to 25% of solid mass and residual TFA can perturb membrane and receptor assay readouts (counterion and orthogonal-testing literature, 2020-12-03).
Pro Tip: Map each remaining uncertainty to the specific orthogonal test that resolves it before you commit assay time.
From a triage shortlist to an assay plan
Work from the shortlist your evidence-state filter produced, then convert each open question into a test and a decision rule.
-
Identity and purity. If the record carries an inferred annotation with no assay context, confirm molecular identity by MS and chromatographic cleanliness by HPLC under a defined method. Resolve the candidate if identity matches and purity meets your specification.
-
Counterion state. Where the record notes a salt form without quantification, measure it. A counterion result that shifts effective dose or perturbs a receptor assay changes the interpretation of everything downstream.
-
Endotoxin. Set the limit from the compendial endotoxin limit, K/M, with K = 5 USP-EU/kg for routes other than intrathecal and K = 0.2 USP-EU/kg for intrathecal, where M is the maximum recommended human dose per kg in a single hour (USP <85> Bacterial Endotoxins Test).
-
Sterility. Demonstrate absence of viable microorganisms under test conditions, which under USP <71> means inoculation into specified media and 14-day incubation (ICH Q2(R2), 2023-11-30).
-
Biological confirmation. Only after the analytical layer is settled should you read a toxicity signal as a property of the candidate rather than of the sample.
Documentation that survives review
Every result above has to be defensible later, and the metadata layer is where that starts. The MHRA’s 2018 GxP data integrity guidance defines ALCOA as attributable, legible, contemporaneous, original, and accurate, with ALCOA+ adding complete, consistent, enduring, and available (MHRA, Guidance on GxP data integrity, 2018-03-09). Provenance fields map onto those attributes directly: a source identifier and curation date make a record attributable and contemporaneous, an assay-context field keeps it complete, and a stable record link keeps it available. Validation expectations run in parallel, since ICH Q2(R2)’s fit-for-purpose standard requires documenting that a procedure suits its intended use, not merely that it ran (ICH Q2(R2), 2026-02-20).
A FAIR metadata layer supports this workflow in a modest, replicable way. It can be used to carry evidence states and provenance fields from triage tooling into CoA-grade records without retyping, and it helps keep the analytical results, the method used, and the record they were matched against in one traceable chain. MOL Changes is built around that documentation and analytical-verification layer, including HPLC/MS and sterility testing and Class 100 sterile manufacture. The practical consequence is that the database never stands in for your testing; it decides which testing you owe, and it supplies the audit trail that shows you ran it.
Benchmarking and Prediction Models Built on Peptide Toxicity Metadata
MAOMAO’s benchmark layer exists so that model performance can be compared on a shared footing rather than on each paper’s private split. It defines 300 endpoint-seed-strategy configurations, 1,500 data partitions, 3,300 representation-specific configurations and 16,500 fold instances, derived from 5 eligible endpoints Syntetyske peptiden × 30 random seeds × 5-fold splitting × 11 representations (MAOMAO, Frontiers in Pharmacology, 2026). The representation side is equally explicit: 10 protein-language-model embedding matrices plus a one-hot representation covering all 71,857 sequences, alongside a descriptor layer of 41 features (MAOMAO, 2026).
The limitation matters as much as the scale. Those predefined partitions use random and stratified splitting and are not homology- or similarity-controlled, so they are not leakage-controlled. A high score on this benchmark tells you a model separates the curated peptide toxicity metadata it was trained on. It does not yet tell you the model transfers to a scaffold it has never seen.
Reading prediction-model claims with the right caveats
Published peptide toxicity predictors report strong numbers, and each number carries its own scope. ToxiPep trained on 3,528 toxic and 3,528 non-toxic peptides filtered at CD-HIT 0.9, tested on 883 plus 883, and reported accuracy 0.850, sensitivity 0.827, specificity 0.873, MCC 0.701 and ToxiPep’s reported 0.92 AUC (2025). ToxGIN’s independent-test AUROC of 0.9172 came with an F1 of 0.8354 and an MCC of 0.6866 (Briefings in Bioinformatics, 2024). PeptiVerse assembled 5,518 toxic and 5,518 non-toxic peptides, but PeptiVerse inherits its toxicity labels from ToxinPred3.0 rather than from assay (2026).
Two caveats follow. First, these are largely vendor-of-own-method results, evaluated by the teams that built the models. Second, applicability domain is bounded by construction: ToxiPep’s CD-HIT 0.9 filtering means its claims hold for sequences similar to its training data, not for arbitrary novel chemistry. PeptiVerse compounds this, because a label inherited from a predictor carries that predictor’s error forward.
What a leakage-controlled benchmark would change
If you are deciding whether to trust a model on a novel scaffold, the split is the claim. Homology- or similarity-controlled partitioning holds near-duplicate sequences out of training, so a test score reflects extrapolation rather than memorisation of close relatives. Random and stratified splits cannot distinguish the two, which is why a 0.92 AUC on a random split and a 0.92 AUC on a homology-controlled split mean different things for your program.
The practical reading: treat published peptide toxicity metadata scores as evidence of internal discrimination, not as a transfer guarantee. Before a model informs a synthesis or testing decision, check whether its evaluation controlled for sequence similarity, and confirm the endpoint and assay context behind its labels. Where that control is absent, the score is a starting hypothesis for orthogonal analytical and biological testing, not a substitute for it.
Common Misconceptions About Peptide Toxicity Databases
The most common error is treating a database label as a safety determination. A label records what was observed, in what system, under what conditions; it does not certify that a sequence is safe to advance.
The first misconception is that a negative label means non-toxic. Negatives are not one thing. A strong negative rests on a documented assay with a stated endpoint and concentration range; a weak negative may reflect an untested sequence, an unreported result, or an assay that could not detect the effect. A triage rule that collapses the two into a single “clean” flag is the failure mode.
The second is that more sources means more confidence. Source count and evidential independence are different quantities. A large aggregate can still trace back to one upstream contributor, so the apparent breadth of support is narrower than the number suggests.
The third is that a model score transfers to a novel sequence. Prediction models carry an applicability domain, and a score computed outside it is an extrapolation, not a measurement.
The fourth is that harmonized data removes the need for orthogonal testing. Harmonization makes records comparable; it does not make them sufficient. Decisions on peptide safety still require qualified professional judgement and orthogonal analytical and biological testing.
Warning: a negative label is not one thing. Strong and weak negatives carry different evidential weight, and a triage rule that treats them as one thing is the failure mode.
Part of why the negative side is structurally thinner is toxicology’s publication bias, measured as a 3.3-fold publication advantage for positive findings, with positive-result prevalence rising from roughly 50% to 90% between 1991 and 2007 (Bero et al., risk-of-bias review, 2016, covering the 1991–2007 window). A literature-derived dataset inherits that asymmetry.
Why non-clinical toxicology attrition makes this a decision, not a curiosity
The stakes are why peptide candidate triage deserves this level of care rather than a quick database lookup. In the four-company attrition audit, non-clinical toxicology accounted for 40% of terminations, the largest single cause, across roughly 605 terminated compounds from AstraZeneca, Eli Lilly, GSK and Pfizer between 2000 and 2010 (Waring et al., Nature Reviews Drug Discovery, 2015).
A pharmacology review reaches the same order of magnitude, putting toxicity at roughly one third of drug-candidate attrition and reporting clinical-development safety failure rates of 35% from phase 1 to submission and 28% from phase 2 to submission (Karnes et al., 2010). The two sources converge on the same conclusion through different samples and methods, which strengthens the signal without making them fully independent estimates.
Getting Started With FAIR Peptide Toxicity Data
The first action is to pull one candidate sequence and read its record end to end before you run anything. That single pass tells you what the record already supports and what it does not.
Step 1: Resolve the sequence and read the record. Resolve the sequence to its deterministic identifier, then read the evidence state and provenance fields attached to it. MAOMAO’s harmonized fields are designed for exactly this pass: a SHA-256 identifier ties the entry to one sequence, and the provenance fields state where the underlying measurement came from. If a field is empty, treat that as information, not as a gap to fill by assumption.
Step 2: Classify each uncertainty and assign a test. Sort what you found into analytical uncertainty and biological uncertainty, then name the orthogonal test that resolves each one. Orthogonal analytical and biological testing means the second method does not share the failure mode of the first, so agreement between them carries weight that a repeat run does not.
Step 3: Record the outcome so it stays reusable. Write the result against the ALCOA+ attributes the MHRA’s ALCOA definition sets out: attributable, legible, contemporaneous, original, accurate, plus complete, consistent, enduring and available. A result recorded this way can be reused in the next program instead of re-derived.
The metadata layer is a starting point, not a substitute for testing. It tells you which experiment to run next; only the experiment tells you what the peptide does.
Frequently Asked Questions
What does MAOMAO stand for, and what does it integrate?
MAOMAO is a curated peptide toxicity resource that brings sequence records, assay context and provenance metadata into one FAIR-aligned layer, so a toxicity label arrives with the conditions that produced it rather than as a bare flag. The integration point matters more than the acronym: sequence, endpoint, assay system and curation date travel together.
Does a negative toxicity label mean a peptide is safe?
Nee. A negative label means the curated record contains no toxicity signal for the tested conditions, which is a statement about the evidence, not about the molecule. Absence of a recorded effect in one assay system does not establish safety, and any development decision still needs qualified toxicological judgement plus orthogonal analytical and biological testing.
How does MAOMAO compare with DBAASP v3?
DBAASP v3, the experimental-record counterpart, covers more than 15,700 peptide entries, including over 14,500 monomeric peptides and roughly 400 homo- and hetero-multimers, with activity and toxicity measurements against 8 target species or cell-type groups (DBAASP v3, released November 2020). The two resources answer different questions: DBAASP v3 aggregates measured records, while a FAIR layer adds the evidence-state and provenance fields that make each record auditable.
Can prediction models like ToxiPep or ToxGIN replace assay testing?
Not as a substitute for experimental confirmation. ToxiPep’s reported 0.92 AUC describes discrimination on its evaluation set, which is a performance statement about a model, not a safety finding about a sequence. Prediction output is best used to rank and prioritise candidates before orthogonal analytical and biological testing, not to close a toxicity question.
What orthogonal testing does a triage shortlist still require?
A shortlist produced from curated peptide toxicity metadata still needs confirmatory analytical work, typically identity and purity assessment by HPLC and mass spectrometry alongside the relevant biological or cell-based assays for the intended endpoint. The database narrows what you test first; it does not remove the testing.
How does FAIR metadata support IND-enabling documentation?
FAIR metadata supports IND-enabling documentation because it keeps each record traceable to its source, method and curation date, which is the same property the MHRA’s ALCOA definition requires of attributable, legible, contemporaneous, original and accurate records. Regulators review the evidence chain, so a well-documented peptide toxicity metadata layer reduces the reconstruction work later.
Conclusion
Curated FAIR peptide toxicity data improves candidate triage, assay design and reporting, and it does not demonstrate safety. That distinction is the whole argument: a provenance-stamped record tells you what was measured, in what system, under what conditions, and where the evidence stops. It does not tell you a peptide is safe, and no database label can carry that weight on its own.
The layered case runs from the data model outward. Evidence states and provenance fields make a record auditable; auditable records make triage defensible; defensible triage produces an assay plan that survives review. Each step depends on the one before it, which is why partial adoption tends to stall at the first ambiguous record.
The resource and its surrounding model ecosystem are still moving. Coverage gaps remain, licensing coverage is incomplete, and benchmarks are not yet leakage-controlled. The MAOMAO release paper reports that 15 of 56 datasets carry explicit licensing information, so verify terms before you build on any table. Treat current prediction-model claims as directional, not settled.
If you are weighing adoption, the useful next step is a scoped conversation about your own triage workflow. Peptide produksje
Disclosure: MOL Changes publishes peptide documentation and analytical verification, including HPLC, MS and sterility testing under Class 100 sterile manufacture. This article is educational and does not recommend or endorse any specific peptide.
Decisions about peptide safety require qualified professional judgement and orthogonal analytical and biological testing, not a database label alone.

