什么是 MAOMAO 以及为什么 FAIR 肽毒性数据很重要?

MAOMAO 代表多源注释组织的元数据感知本体, 本体引导的资源,协调从公共数据库中提取的肽毒性信息, 文献衍生资源和精选数据集 (MAOMAO离型纸, 2026). 它不是具有单个上游的单个数据库. 它的第一个版本集成了 56 来源, 跨越公共数据库, 文献来源的资源, 补充数据集和手动管理的存储库 (同一张纸), 并将它们合并为 71,857 跨越八个毒性相关终点的独特肽序列. 核心层持有 574,856 这些序列的协调肽端点组合. 6 月份期间 Google Scholar 和 PubMed 上的搜索量 2025 和九月 2026, 因此该版本反映了截至 9 月份的可用资源 2026. 此类数字因来源和管理日期而异, 并且版本中的单个上游记录仅承载该上游的权重.
FAIR 肽毒性数据的含义比“组织良好的数据”更窄。它意味着带有持久标识符的毒性记录, 丰富的来源元数据和规定的许可证, 这样读者就可以追踪声明的来源以及在什么条件下可以重复使用它.
类比: 记录是实验室笔记本条目, 不是判决
像对待带有出处标记的笔记本条目一样对待毒性注释. 它记录了谁观察到了效果, 什么时候, 来自哪个来源, 在哪个生物体上, 以及根据哪些证据表明. 它并不证明序列是安全的. 这种区别影响了本指南中的每个下游决策, 因为记录会告诉您所主张的内容以及记录的情况, 不是你对自己的候选人所证明的.
FAIR 对毒性记录的实际要求是什么
缩写隐藏了操作要求. 这 原创公平原则 (威尔金森等人。, 科学数据, 2016) 和 GO FAIR 概述 指定数据携带持久标识符 (F1), 元数据丰富 (F2) 并包括它们所描述的数据的标识符 (F3), 详细出处 (R1.2), 该记录遵循领域相关的社区标准 (R1.3), 并且有一个明确的, 注明可访问使用许可证 (R1.1).

MAOMAO统一领域是具体落实: 源标识符, 文献参考, 有机体信息, 注释出处, 检索日期, 文件格式和版本信息(如果有), 每个标准化序列都由大写的确定性 SHA-256 标识符键入, 无空格序列. 文件以 CSV 形式发送, JSON 和 YAML.
要点: 精心策划的毒性记录是带有出处标记的声明, 不是安全判决.
为什么策划的肽毒性元数据会改变候选分类

Curated peptide toxicity metadata changes peptide candidate triage because it lets you filter on the strength of the underlying evidence rather than on the presence of a label. A database that stores only “toxic” or “non-toxic” collapses very different observations into one field. MAOMAO separates them. Its evidence states distinguish positive, negative, ambiguous, unlabeled and no-information records, and the negative state splits further: strong negatives were explicitly supported by experimental evidence, while weak negatives originated from randomly selected negative sets or lacked confirmation of the absence of toxicity annotations (MAOMAO离型纸, 2026).
The distribution shows why that distinction carries weight. Across the release, MAOMAO reports 31,153 positive, 91,506 negative, 15,717 ambiguous and 5,044 unlabeled states (MAOMAO’s own documentation of its evidence states, 检索到的 2026-07-30). The negative bucket is nearly three times the positive bucket, so a triage rule that treats every negative as clearance is not filtering on evidence. It is filtering on volume.
|
Evidence state |
Count |
What it supports in triage |
|---|---|---|
|
多肽合成 积极的 |
31,153 |
Flag for follow-up; check endpoint and assay context |
|
Negative |
91,506 |
Depends on strong vs weak origin |
|
Ambiguous |
15,717 |
Cannot resolve a go/no-go on its own |
|
Unlabeled |
5,044 |
No usable signal |
阅读一条记录并确定它证明了什么
Take a record carrying an inferred annotation with no assay context attached. The label says negative. The provenance says the call was inferred rather than measured, and no endpoint, concentration or assay system is recorded alongside it. At that point the record does not answer “is this sequence safe to advance.” It answers “has anyone published a sequence-level toxicity observation for this sequence,” and the answer is no.
That is the decision point. A triage rule built on the presence of a label passes the candidate through. A rule built on evidence state routes it to the ambiguous or unlabeled path, where it needs either a literature check or an assay before it earns a place on a shortlist. The second rule costs more per candidate and produces a shortlist you can defend.
The same logic applies when two sources disagree on the endpoint label for one sequence. The curated record does not pick a winner silently. It holds both annotations with their provenance, and the disagreement itself becomes the signal that the sequence needs orthogonal analytical and biological testing rather than a database verdict.
证据耗尽的地方
Endpoint coverage is uneven, and the gap is structural rather than incidental. MAOMAO states that coverage depends on retrievable sequence-level evidence, which produces class imbalance and limited representation of Cytolysis, Embryotoxic and Ichthyotoxic endpoints (MAOMAO离型纸, 2026). The counts make the thinness concrete: Embryotoxic carries 2 positive states and Ichthyotoxic 5, while Cytolysis holds 343 states in total. Compare that with Toxic root at 57,080 states and 13,824 positives, or Hemolytic at 35,960 states and 4,714 positives.
For a triage rule, this means absence of an annotation in a thin endpoint is not evidence of absence. It means the endpoint was rarely recorded. A shortlist filtered on Embryotoxic or Ichthyotoxic data is filtered on almost nothing, and the honest move is to say so before the candidate reaches a program decision.
Licensing is the second boundary. Annotation content and last-update dates are available for 55 的 56 来源, but explicit licensing covers only 15 的 56 (MAOMAO’s own documentation of its evidence states, 检索到的 2026-07-30). Reuse terms for the remainder need checking at the source, not assumed from the aggregate.
要点: A negative is not one thing. Triage on evidence state and provenance, 店铺 not on label presence, and treat thin endpoints as unmeasured rather than clean.
肽毒性证据状态和来源字段, 逐个领域
A peptide toxicity record is auditable when two things are explicit: what the evidence state is, and where the record came from. MAOMAO structures both. The ontology has a root class, Toxic, with primary classes Cytotoxic, Neurotoxic, Embryotoxic, and Ichthyotoxic, and Cytotoxic carries subclasses Hemolytic, Cytolysis, and Anti mammalian cells (MAOMAO release paper, 2026). Read branch totals with care: the root Toxic overlaps subclass counts, so summing the branches will overstate the root, a limitation the paper’s methodology note states directly.
Identity is handled by hashing the normalized uppercase sequence with SHA-256, which gives the same sequence the same key across layers and sources. Records ship as CSV, JSON, or YAML.
实践中的证据状态词汇
Each state answers a different question for a program. 积极的: is there direct or hierarchy-supported toxicity evidence? Negative: is there direct, endpoint-specific non-toxic evidence? Ambiguous: are annotations conflicting or unresolved? Unlabeled: was the sequence retained without an explicit class assignment? No information: is there simply no evidence for this endpoint? (MAOMAO release paper, 2026). Treat these as documentation of what the source said, not as a safety determination.
使记录可审计的出处字段
MAOMAO’s harmonized fields let a reviewer check a claim without opening the original source. Source identifier and literature reference satisfy Findable and traceable attribution. Organism and annotation provenance satisfy Reusable context. Retrieval date and version satisfy Accessible currency. File format supports Interoperable exchange. 关于
Licensing is the honest gap. The paper states no unified dataset license in its main text, reports source-level licensing coverage of 15 的 56 来源, and places licensing information and versioning records in the Documentation Layer (MAOMAO release paper, 2026). Check source-level terms before reuse in a regulated workflow.
FAIR 肽毒性数据如何融入检测设计和报告

Curated peptide toxicity metadata tells you which orthogonal confirmation to run next; it does not tell you a candidate is safe. That boundary is the whole point of the layer. A database record summarizes what someone else observed under stated conditions, while demonstrated safety rests on orthogonal analytical and biological testing performed on your material: 高效液相色谱纯度, 女士身份, counterion state, endotoxin by LAL, 不育, and the biological assays your program requires. Each test answers a narrow question. LAL measures endotoxin burden only, not sterility or chemical purity, and counterion mass is not trivial, since the TFA salt can represent roughly 5% 到 25% of solid mass and residual TFA can perturb membrane and receptor assay readouts (counterion and orthogonal-testing literature, 2020-12-03).
对于小费: Map each remaining uncertainty to the specific orthogonal test that resolves it before you commit assay time.
从分类候选名单到检测计划
Work from the shortlist your evidence-state filter produced, then convert each open question into a test and a decision rule.
-
身份和纯度. If the record carries an inferred annotation with no assay context, confirm molecular identity by MS and chromatographic cleanliness by HPLC under a defined method. Resolve the candidate if identity matches and purity meets your specification.
-
Counterion state. Where the record notes a salt form without quantification, measure it. A counterion result that shifts effective dose or perturbs a receptor assay changes the interpretation of everything downstream.
-
内毒素. Set the limit from the compendial endotoxin limit, K/M, 与 K = 5 USP-EU/kg for routes other than intrathecal and K = 0.2 USP-EU/kg 鞘内注射, where M is the maximum recommended human dose per kg in a single hour (美国药典 <85> 细菌内毒素检查).
-
不育. Demonstrate absence of viable microorganisms under test conditions, which under USP <71> means inoculation into specified media and 14-day incubation (我Q2(R2), 2023-11-30).
-
Biological confirmation. Only after the analytical layer is settled should you read a toxicity signal as a property of the candidate rather than of the sample.
通过审查的文档
Every result above has to be defensible later, and the metadata layer is where that starts. The MHRA’s 2018 GxP data integrity guidance defines ALCOA as attributable, 清晰易读, 同时期的, 原来的, and accurate, with ALCOA+ adding complete, 持续的, 持久, 并可用 (MHRA, Guidance on GxP data integrity, 2018-03-09). Provenance fields map onto those attributes directly: a source identifier and curation date make a record attributable and contemporaneous, an assay-context field keeps it complete, and a stable record link keeps it available. Validation expectations run in parallel, since ICH Q2(R2)’s fit-for-purpose standard requires documenting that a procedure suits its intended use, not merely that it ran (我Q2(R2), 2026-02-20).
A FAIR metadata layer supports this workflow in a modest, replicable way. It can be used to carry evidence states and provenance fields from triage tooling into CoA-grade records without retyping, and it helps keep the analytical results, the method used, and the record they were matched against in one traceable chain. MOL Changes is built around that documentation and analytical-verification layer, including HPLC/MS and sterility testing and Class 100 sterile manufacture. The practical consequence is that the database never stands in for your testing; it decides which testing you owe, and it supplies the audit trail that shows you ran it.
基于肽毒性元数据构建的基准测试和预测模型
MAOMAO’s benchmark layer exists so that model performance can be compared on a shared footing rather than on each paper’s private split. It defines 300 endpoint-seed-strategy configurations, 1,500 data partitions, 3,300 representation-specific configurations and 16,500 fold instances, 源自 5 eligible endpoints 合成肽 × 30 random seeds × 5-fold splitting × 11 representations (MAOMAO, 药理学前沿, 2026). The representation side is equally explicit: 10 protein-language-model embedding matrices plus a one-hot representation covering all 71,857 序列, alongside a descriptor layer of 41 features (MAOMAO, 2026).
The limitation matters as much as the scale. Those predefined partitions use random and stratified splitting and are not homology- or similarity-controlled, so they are not leakage-controlled. A high score on this benchmark tells you a model separates the curated peptide toxicity metadata it was trained on. It does not yet tell you the model transfers to a scaffold it has never seen.
阅读具有正确警告的预测模型声明
Published peptide toxicity predictors report strong numbers, and each number carries its own scope. ToxiPep trained on 3,528 toxic and 3,528 non-toxic peptides filtered at CD-HIT 0.9, tested on 883 plus 883, and reported accuracy 0.850, 灵敏度 0.827, 特异性 0.873, MCC 0.701 和 ToxiPep’s reported 0.92 曲线下面积 (2025). ToxGIN’s independent-test AUROC 的 0.9172 came with an F1 of 0.8354 and an MCC of 0.6866 (Briefings in Bioinformatics, 2024). PeptiVerse assembled 5,518 toxic and 5,518 non-toxic peptides, but PeptiVerse inherits its toxicity labels from ToxinPred3.0 rather than from assay (2026).
Two caveats follow. 第一的, these are largely vendor-of-own-method results, evaluated by the teams that built the models. 第二, applicability domain is bounded by construction: ToxiPep’s CD-HIT 0.9 filtering means its claims hold for sequences similar to its training data, not for arbitrary novel chemistry. PeptiVerse compounds this, because a label inherited from a predictor carries that predictor’s error forward.
泄漏控制基准会改变什么
If you are deciding whether to trust a model on a novel scaffold, the split is the claim. Homology- or similarity-controlled partitioning holds near-duplicate sequences out of training, so a test score reflects extrapolation rather than memorisation of close relatives. Random and stratified splits cannot distinguish the two, which is why a 0.92 AUC on a random split and a 0.92 AUC on a homology-controlled split mean different things for your program.
The practical reading: treat published peptide toxicity metadata scores as evidence of internal discrimination, not as a transfer guarantee. Before a model informs a synthesis or testing decision, check whether its evaluation controlled for sequence similarity, and confirm the endpoint and assay context behind its labels. Where that control is absent, the score is a starting hypothesis for orthogonal analytical and biological testing, not a substitute for it.
关于肽毒性数据库的常见误解
The most common error is treating a database label as a safety determination. A label records what was observed, in what system, under what conditions; it does not certify that a sequence is safe to advance.
The first misconception is that a negative label means non-toxic. Negatives are not one thing. A strong negative rests on a documented assay with a stated endpoint and concentration range; a weak negative may reflect an untested sequence, an unreported result, or an assay that could not detect the effect. A triage rule that collapses the two into a single “clean” flag is the failure mode.
The second is that more sources means more confidence. Source count and evidential independence are different quantities. A large aggregate can still trace back to one upstream contributor, so the apparent breadth of support is narrower than the number suggests.
The third is that a model score transfers to a novel sequence. Prediction models carry an applicability domain, and a score computed outside it is an extrapolation, not a measurement.
The fourth is that harmonized data removes the need for orthogonal testing. Harmonization makes records comparable; it does not make them sufficient. Decisions on peptide safety still require qualified professional judgement and orthogonal analytical and biological testing.
警告: a negative label is not one thing. Strong and weak negatives carry different evidential weight, and a triage rule that treats them as one thing is the failure mode.
Part of why the negative side is structurally thinner is toxicology’s publication bias, measured as a 3.3-fold publication advantage for positive findings, with positive-result prevalence rising from roughly 50% 到 90% between 1991 和 2007 (Bero et al., risk-of-bias review, 2016, covering the 1991–2007 window). A literature-derived dataset inherits that asymmetry.
为什么非临床毒理学消耗会做出这样的决定, 不是好奇心
The stakes are why peptide candidate triage deserves this level of care rather than a quick database lookup. In the four-company attrition audit, non-clinical toxicology accounted for 40% of terminations, the largest single cause, 大致穿过 605 terminated compounds from AstraZeneca, Eli Lilly, GSK and Pfizer between 2000 和 2010 (Waring et al., Nature Reviews Drug Discovery, 2015).
A pharmacology review reaches the same order of magnitude, putting toxicity at roughly one third of drug-candidate attrition and reporting clinical-development safety failure rates of 35% from phase 1 to submission and 28% from phase 2 to submission (Karnes et al., 2010). The two sources converge on the same conclusion through different samples and methods, which strengthens the signal without making them fully independent estimates.
FAIR 肽毒性数据入门
The first action is to pull one candidate sequence and read its record end to end before you run anything. That single pass tells you what the record already supports and what it does not.
步 1: Resolve the sequence and read the record. Resolve the sequence to its deterministic identifier, then read the evidence state and provenance fields attached to it. MAOMAO’s harmonized fields are designed for exactly this pass: a SHA-256 identifier ties the entry to one sequence, and the provenance fields state where the underlying measurement came from. If a field is empty, treat that as information, not as a gap to fill by assumption.
步 2: Classify each uncertainty and assign a test. Sort what you found into analytical uncertainty and biological uncertainty, then name the orthogonal test that resolves each one. Orthogonal analytical and biological testing means the second method does not share the failure mode of the first, so agreement between them carries weight that a repeat run does not.
步 3: Record the outcome so it stays reusable. Write the result against the ALCOA+ attributes the MHRA’s ALCOA definition sets out: 可归因的, 清晰易读, 同时期的, 原来的, 准确的, 加完整, 持续的, enduring and available. A result recorded this way can be reused in the next program instead of re-derived.
The metadata layer is a starting point, not a substitute for testing. It tells you which experiment to run next; only the experiment tells you what the peptide does.
常见问题解答
MAOMAO 代表什么, 它集成了什么?
MAOMAO is a curated peptide toxicity resource that brings sequence records, assay context and provenance metadata into one FAIR-aligned layer, so a toxicity label arrives with the conditions that produced it rather than as a bare flag. The integration point matters more than the acronym: 顺序, endpoint, assay system and curation date travel together.
毒性标签阴性是否意味着肽是安全的?
不. A negative label means the curated record contains no toxicity signal for the tested conditions, which is a statement about the evidence, not about the molecule. Absence of a recorded effect in one assay system does not establish safety, and any development decision still needs qualified toxicological judgement plus orthogonal analytical and biological testing.
MAOMAO 与 DBAASP v3 相比如何?
DBAASP v3, the experimental-record counterpart, covers more than 15,700 peptide entries, including over 14,500 monomeric peptides and roughly 400 homo- and hetero-multimers, with activity and toxicity measurements against 8 target species or cell-type groups (DBAASP v3, released November 2020). The two resources answer different questions: DBAASP v3 aggregates measured records, while a FAIR layer adds the evidence-state and provenance fields that make each record auditable.
ToxiPep 或 ToxGIN 等预测模型能否取代分析测试?
Not as a substitute for experimental confirmation. ToxiPep’s reported 0.92 AUC describes discrimination on its evaluation set, which is a performance statement about a model, not a safety finding about a sequence. Prediction output is best used to rank and prioritise candidates before orthogonal analytical and biological testing, not to close a toxicity question.
分类候选名单还需要哪些正交测试?
A shortlist produced from curated peptide toxicity metadata still needs confirmatory analytical work, typically identity and purity assessment by HPLC and mass spectrometry alongside the relevant biological or cell-based assays for the intended endpoint. The database narrows what you test first; it does not remove the testing.
FAIR 元数据如何支持 IND 文档?
FAIR metadata supports IND-enabling documentation because it keeps each record traceable to its source, method and curation date, which is the same property the MHRA’s ALCOA definition requires of attributable, 清晰易读, 同时期的, original and accurate records. Regulators review the evidence chain, so a well-documented peptide toxicity metadata layer reduces the reconstruction work later.
结论
Curated FAIR peptide toxicity data improves candidate triage, assay design and reporting, and it does not demonstrate safety. That distinction is the whole argument: a provenance-stamped record tells you what was measured, in what system, under what conditions, and where the evidence stops. It does not tell you a peptide is safe, and no database label can carry that weight on its own.
The layered case runs from the data model outward. Evidence states and provenance fields make a record auditable; auditable records make triage defensible; defensible triage produces an assay plan that survives review. Each step depends on the one before it, which is why partial adoption tends to stall at the first ambiguous record.
The resource and its surrounding model ecosystem are still moving. Coverage gaps remain, licensing coverage is incomplete, and benchmarks are not yet leakage-controlled. The MAOMAO release paper reports that 15 的 56 datasets carry explicit licensing information, so verify terms before you build on any table. Treat current prediction-model claims as directional, not settled.
If you are weighing adoption, the useful next step is a scoped conversation about your own triage workflow. 多肽生产
披露: MOL Changes publishes peptide documentation and analytical verification, including HPLC, MS and sterility testing under Class 100 sterile manufacture. This article is educational and does not recommend or endorse any specific peptide.
Decisions about peptide safety require qualified professional judgement and orthogonal analytical and biological testing, not a database label alone.

