肽折叠模型和聚集实际预测什么
折叠模型是一个地区的天气预报, 不是对您所在街道的测量. 这种区别决定了肽聚集预测的实际限制是什么, 在第一次模拟运行之前值得说明.

多肽合成 三个术语构成本文的其余部分. 一个 力场 是为系统中每个原子或珠子分配能量的方程和参数集. 一个 粗粒度模型 用单个相互作用位点取代原子团, 交易化学细节以获得可达到的时间范围. 一个 构象系综 是肽访问的结构的集合, 按概率加权, 而不是一种获胜形状.
简化和粗粒度的模型可以很好地回答集成级问题. 他们排名聚集趋势, 追踪自由能趋势, 将富含 β 的物质与受保护状态分开, 读取疏水图案, 并估计相对自我关联风险. 它们对结构特异性不可靠或沉默: 精确的低聚物尺寸分布, 原纤维与无定形形态, 和序列特异性多态性不属于这些方法解决的范围. 报告的聚合行为也依赖于力场, 对于同一肽系统,不同的力场可能会返回矛盾的结果, 如记录在 2025 蛋白质聚集体的结构表征和随附的力场可转移性评估指南.
要点: 模型输出是一个排名, 不是一个利率. 它告诉您哪个序列或条件更容易聚合, 不知道有多快, 变成什么形态, 或者您的肽实际组装的寡聚体大小是多少.
合成肽 src=”https://molchanges.com/wp-content/uploads/2026/09/pub_20260922_044025_896_59a6098c9afc48d9bc0bd7b9f580aa27.png”>
为什么相同的肽会得到不同的答案: 力场和采样限制
同一序列的两次模拟可能会因与肽无关的原因而不一致. 第一个是力场. 力场可转移性评估发现,同一系统在一个参数集下可以解析为富含 β 的聚集体,而在另一个参数集下则可以解析为很大程度上无序。, 以采样窗口作为混杂因素: 在一种情况下,在几百纳秒内富含β, 粗略地之后才变得无序 2 另一个微秒 (PMC8120800). 这是一个上游基准, 不是一个趋同的研究机构, 应该这样读.
第二个限制是采样和系统规模. 一项公正的全原子研究收集了 75 微秒的淀粉样蛋白肽自组装过程 12 立方盒中的肽分子, 达到十二聚体的低聚物, N=4 以上低聚物的首次传代时间仍然分散在 10 跨越独立轨迹的纳秒和数百纳秒 (科学报告, 2016).
这两个限制都不是您可以忽略的错误. 力场不一致是单个模拟不能证明聚合风险的原因, 这就是为什么肽聚集预测限制属于报告的方法部分, 不在脚注中.
仿真与基准之间的专注度和时间尺度差距
模拟聚合几乎总是比它旨在告知的实验集中得多. 一个 2025 混合粗粒度研究,大致模拟 27 mM,当实验接近时 1 mM 是一个有用的说明 (分子, 2025). 我找不到那个单一的上游结果, 因此将其视为间隙的示例而不是校准常数. 多肽生产
差距很重要,因为物理主导的浓度发生了变化. 在 27 毫米, 碰撞足够频繁,以至于模拟系统在微秒轨迹可以达到的时间尺度上聚合. 在 1 毫米, 相同的序列可能稳定几个小时. 因此,微秒采样不会映射到真正的聚合分析所占用的数小时到数天的窗口中, 两个间隙复合: 高浓度会加速你想要计时的事件.
对于小费: 远高于您工作浓度的模拟运行可以对序列进行相互排名, 但它无法为您提供绝对聚集率或预测您工作浓度下的配方行为.
What the high-concentration result licenses is a relative claim: sequence A aggregates faster than sequence B under identical simulated conditions. What it does not license is an absolute one: that A will aggregate in your buffer, at your concentration, within your stability window. Keep those two claims separate and the simulation stays useful.
当模拟足够时: 筛选易于聚集的序列
Simulation is enough when the decision is a ranking among candidates, not a prediction of behavior. In-silico triage works as a funnel: screen the full candidate set with fast sequence-level predictors, keep only candidates acceptable across several orthogonal tools because different algorithms emphasize different features, map flagged hotspots onto sequence and structure, then synthesize and test only the top-ranked constructs. 作为 a review of in-silico aggregation algorithms 说它, these tools rank and flag rather than predict with certainty, and they do not replace HIC, 美国证券交易委员会, 动态LS, HPLC or stability assays.
That boundary is not theoretical. 一个 2026 Nature Chemistry analysis of 539 peptide sequences found whole-sequence XGBoost predicted on-resin aggregation at 58.0% ± 3.5% 准确性, barely above the 57.7% ± 3.3% it reached on residue-shuffled sequences, while a composition-vector representation scored 59.5% ± 1.9% and the earlier Mohapatra model reached 60%. Shuffling preserved behavior: 19 的 20 shuffled aggregating peptides stayed aggregating, 和 14 的 20 shuffled non-aggregating sequences stayed non-aggregating. 一个 directed-evolution platform that screened a 1.3-million-mutant library found the same disagreement from the other side: three algorithms flagged 26 residues as aggregation-suppressible, but covered only 8 的 12 evolution hit residues, and just 3 residues were flagged by all three.
Use aggregation-prone sequence screening to order your work, then let the bench settle it.
|
Model output |
What it licenses you to conclude |
What still needs bench data |
|---|---|---|
|
Aggregation tendency |
Relative ranking within one candidate set |
Absolute aggregation propensity |
|
β-sheet propensity |
Which segments to mutate or protect |
Whether substitution changes crude purity |
|
Oligomer size distribution |
Whether oligomers are plausible |
Actual size distribution (美国证券交易委员会, 动态LS) |
|
Morphology |
Which morphologies to expect |
Fibril versus amorphous form (cryo-TEM) |
|
Polymorphism |
That multiple forms may coexist |
Which form your batch adopts |
预测器在相反方向上失败的地方
AlphaFold-class tools and sequence-only aggregation predictors miss the same risk from opposite sides, which is why neither one alone settles whether a variant will aggregate. 一个 2025 review of computational aggregation prediction found that AlphaFold-class models return one dominant static structure: they do not sample conformational ensembles or partially folded intermediates, they carry no pH, 离子强度, 温度, solvent or membrane context, and their confidence scores measure structural consistency rather than aggregation propensity. Sequence-only peptide predictors invert the error. Many are hexapeptide-centred and cannot tell whether a motif will end up buried, exposed, or membrane-associated, so they over-predict aggregation in folded proteins while performing better on isolated peptides and intrinsically disordered regions.
The practical consequence is that the two failure modes are complementary rather than redundant. A high-confidence fold says nothing about a conditionally exposed sticky segment, and a flagged hexapeptide says nothing about whether that segment is reachable in the folded state. Treating either output as a verdict on peptide aggregation prediction limits the screen to one class of error, so orthogonal methods that cover both the structural and the sequence view are the defensible default.
当您需要定制合成时, 表征, 和化验
Simulation answers relative questions. Bench work answers absolute ones: a rate, a morphology, a formulation sensitivity. When your decision depends on any of those, the model output is a starting hypothesis, not a result.
Each method answers one question and carries its own blind spot. Circular dichroism reports global secondary-structure content, but it is an ensemble average that cannot localize where a change occurred, and it is vulnerable to scattering artifacts; usable concentrations run from 0.05 到 50 mg/mL when concentration in mg/mL times pathlength in mm stays near 0.1 到 0.2 (一个 2025 practitioner’s guide to aggregate characterization, 2025). FTIR reads the amide I region and tolerates concentrated or aggregated samples, yet it is still global and its band assignments remain ambiguous.
DLS gives hydrodynamic size and polydispersity, but it cannot separate monomer from dimer and it is biased toward larger particles because scattering scales with size; practical ranges sit near 0.2 到 50 毫克/毫升, and for large aggregates the single-scattering ceiling can fall to about 0.5 毫克/毫升 (an application guide to DLS for peptide aggregation, 2026). SEC-MALS adds fractionation and absolute molecular weight, 覆盖 200 Da to 1 GDa and radii of gyration from 10 到 500 nm on a DAWN, with a high-concentration option up to 180 毫克/毫升 (the instrument vendor’s published SEC-MALS specifications, 检索到的 2026), though fractionation can dilute or perturb weak reversible assemblies.
ThT tracks amyloid kinetics with sigmoidal traces, 通常在 10 到 20 微米, 和 20 到 50 µM giving the highest signal and 50 µM used for endpoint quantification, but it is insensitive to native protein and to many oligomeric or early intermediates (the standard ThT concentration study, 2017). Cryo-TEM shows morphology directly while imaging only a small fraction of the population, and SAXS returns a low-resolution, ensemble-averaged solution shape that cannot uniquely resolve heterogeneous mixtures (a practical survey of the analytical techniques used to characterize protein and peptide aggregates, 2021).
|
方法 |
Aggregation question it answers |
Detection limit or interference risk |
|---|---|---|
|
光盘 |
Global secondary-structure content |
Ensemble average, cannot localize change; scattering artifacts; 0.05–50 mg/mL with concentration × pathlength ≈ 0.1–0.2 |
|
FTIR |
Global secondary structure, amide I |
Still global; band assignment ambiguous; tolerates concentrated samples |
|
动态LS |
Hydrodynamic size and polydispersity |
Cannot separate monomer from dimer; biased toward larger particles; ≈0.2–50 mg/mL, single-scattering ceiling ≈0.5 mg/mL for large aggregates |
|
SEC-MALS |
Monomer/oligomer populations, absolute molecular weight |
Fractionation can dilute or perturb weak reversible assemblies; less informative about shape |
|
ThT |
Amyloid fibril formation and kinetics |
Insensitive to native protein and to many oligomeric or early intermediates; 10–20 µM typical |
|
Cryo-TEM |
Morphology |
Images only a small fraction of the population |
|
SAXS |
Low-resolution solution shape |
Ensemble-averaged; cannot uniquely resolve heterogeneous mixtures |
The synthesis side carries its own uncertainty. Difficult-sequence SPPS is not a rare edge case: 穿过 539 peptide sequences, 49.9% showed on-resin aggregation, defined as deprotection peak broadening above 20% versus the first coupling, typically beginning 5 到 15 residues from the resin anchor (一个 2026 Nature Chemistry analysis of 539 peptide sequences, 2026). The same analysis found that pseudoproline incorporation lifted crude purity from 23% 到 69% for hGH and from 17% 到 75% for GB1, so aggregation-control strategies pay off as purity rather than as prediction.
警告: Report the conditions your assay ran under: 温度, 酸碱度, 离子强度, peptide concentration, 孵化时间, and agitation history. Without them, a morphology or a rate cannot be compared to any other dataset, including your own earlier runs.
A neutral, replicable example of closing a gap the model left open: when a folding model flags a sequence as aggregation-prone but cannot say whether the aggregate is fibrillar or amorphous, the answer comes from the bench. MOL Changes supports that handoff through custom peptide synthesis for aggregation studies, combining SPPS, RP-HPLC purification to 95% 或更高, and ESI-MS identity confirmation so the material entering CD, ThT, or SEC-MALS is the sequence you modeled rather than a mixture of truncation and deletion products. Purity specifications, 抗衡离子形式, and endotoxin limits should be stated per lot, because each one shifts the readout a downstream assay produces.
设计检测方法,使结果有意义
An aggregation assay is a factorial matrix, not a single measurement. Vary concentration, 酸碱度, 离子强度, 温度, and time together, because the aggregate a model predicts is the one that forms under a specific combination of those variables. Run buffer-only, protein-only, and dye or scattering controls alongside the sample, and take enough time points to separate onset from progression. A single endpoint reading cannot distinguish the two.
服务 The distinction that decides what you can claim is kinetic versus thermodynamic control. Under kinetic control, the observed aggregate is the fastest-forming species under those conditions. Under thermodynamic control, it is the most stable or equilibrium species after enough time for reversibility and rearrangement. A transthyretin aggregation protocol study lays out how these regimes are separated experimentally and what a protocol must report for the comparison to hold.
That reporting discipline is why so many published aggregation comparisons are not comparable: different matrices, different time windows, different controls. Get this step right and a model’s ranking becomes a defensible absolute statement about your peptide.
关于肽折叠模型和聚集的常见误解
A confidence score is not an aggregation score. pLDDT and similar metrics describe how well a model reproduces its own predicted fold, not whether that fold will self-associate in solution. The fix is to treat confidence as a filter on structural plausibility and to source aggregation risk from dedicated predictors or, better, from measured behavior.
One force field is not ground truth. The same sequence can be predicted aggregation-prone under one parameter set and benign under another, which is why single-run outputs should be reported as method-dependent rather than as the system’s property. Run at least two force fields or sampling protocols and treat disagreement as the signal.
Simulated concentration is not working concentration. The gap between simulation-scale and bench-scale concentration is orders of magnitude, so a model that behaves well at high simulated concentration says little about a 1 mM experimental preparation. Match the model’s regime to the experiment before drawing conclusions.
Sequence-only predictors are not built for folded proteins. These tools over-predict aggregation risk when applied to sequences that fold into stable globular structures, because buried hydrophobic stretches are not exposed in the native state. Confirm the folded context before accepting a high-risk call.
One assay is not the whole picture. 光盘, 动态LS, SEC-MALS, ThT, cryo-TEM and SAXS each answer a different question and each has a blind spot: ThT misses non-fibrillar aggregates, DLS struggles to resolve mixtures, and CD reports average secondary structure rather than species distribution. Pair orthogonal methods before declaring a formulation clean.
要点: If the decision needs an absolute number, a morphology, or a formulation answer, the model is not the instrument.
成功是什么样的: 一个站得住脚的聚合结论
If the workflow ran correctly, you can now state five things without hedging. 第一的, which candidates you deprioritized and on what basis: a specific model output, not a general impression that they looked risky. 第二, which single construct advanced to synthesis. 第三, what identity and purity confirmation showed. 第四, what the secondary-structure readout reported. Fifth, what the aggregation assay measured under a fully specified condition set.
Those five markers are concrete and auditable. Purity is reported as an RP-HPLC figure against a stated threshold, commonly 95% or higher for assay-grade material. Identity is confirmed by ESI-MS against the calculated mass. Secondary structure comes from a CD or FTIR readout. The assay contributes a time course with onset separated from progression, so you can distinguish a construct that aggregates immediately from one that degrades over hours.
Once that baseline condition set is fixed, the obvious stretch goal is a formulation-sensitivity or short stability study on the same construct. You already have the reference point; changing one variable at a time from there is cheap compared with re-establishing the baseline.
常见问题解答
我可以单独使用 AlphaFold 对易于聚合的变体进行排名吗?
不. AlphaFold predicts a single static structure, so it returns no conformational ensemble and no environmental context such as pH, 离子强度, or concentration. Aggregation depends on transient exposed patches and on the population of partially unfolded states, neither of which a single model captures. Use AlphaFold as one input, then rank candidates with an explicit aggregation predictor or a short simulation ensemble.
为什么同一肽的两次模拟结果不一致?
Force fields encode different torsional preferences and different treatment of solvation, so the same sequence can populate different secondary-structure propensities under each. Sampling adds a second source of divergence: a run that never escapes its starting basin reports the starting basin, not the accessible ensemble. Before comparing two results, check which force field and which effective sampling window each used.
聚合测定需要多长时间, 以及什么决定了持续时间?
Duration follows the question. A kinetic readout such as thioflavin T fluorescence resolves nucleation and elongation over hours to days, while a thermodynamic endpoint such as a solubility 店铺 or sedimentation measurement needs the system to reach equilibrium, which can take longer. Set the window from the process you are trying to observe, not from instrument convenience.
如果我只能运行一个方法,我应该使用哪一种方法?
There is no single method. Map the question to the technique: RP-HPLC and ESI-MS for identity and purity, circular dichroism or FTIR for secondary structure, DLS or SEC-MALS for oligomer size, and cryo-TEM for morphology. One method answers one question, so pick the one that tests the specific failure mode you suspect.
高浓度模拟聚集是否仍然有用?
Yes for ranking, no for absolute rates. Simulated systems often sit far above the experimental working concentration, so the ordering of variants can transfer while the predicted timescale does not. Treat the output as a relative ranking and confirm the top candidates at bench concentration.
对于一个困难的序列我应该期望什么纯度?
Aggregation-prone sequences are harder to make, and on-resin aggregation during chain assembly is a common cause of low crude purity. Pseudoproline dipeptide building blocks disrupt that on-resin aggregation and improve the purity of difficult sequences, so the achievable specification depends on the sequence and on the synthesis strategy chosen for it.
结论
You can now separate the questions a simplified peptide folding models and aggregation workflow can answer from the ones that only synthesis, 表征, and assay can settle. Models rank and flag candidate sequences; they do not predict aggregation with certainty, and they do not replace HIC, 美国证券交易委员会, 动态LS, 高效液相色谱法, or stability studies, as a review of in-silico aggregation algorithms sets out. Treat a prediction as a hypothesis to test, not a result to report. 关于
If the next step is generating that evidence, MOL Changes supports custom synthesis, 纯化, and characterization for aggregation-prone sequences, from milligram screening lots to kilogram scale. You can talk to an expert about a characterization consultation and scope the assays your sequence actually needs.
披露: MOL Changes provides peptide synthesis and characterization services, so it has a commercial interest in the quality standards discussed here.
