Academic Jobs - Home of Higher Ed Logo

Modelos de predicción de estructura molecular en el laboratorio

Publicar una historia
1056Opinión
A scientist in a white lab coat using a laptop in a laboratory
Photo by ThisisEngineering on Unsplash

The first time a laboratory predicts a protein structure on a laptop, the response is usually wary optimism. A graduate student pastes an amino acid sequence into a web server, waits a few minutes to an hour, and receives a three-dimensional model with color-coded confidence scores. The model looks like a real structure: alpha helices, beta sheets, and loops that tangle into a plausible active site. It also carries a number that separates the parts to trust from the parts to ignore. That number, the predicted local distance difference test or pLDDT, has become one of the most read values in structural biology. Five years ago, most bench scientists had not heard of it.

Computational methods for predicting molecular structure are not new. Homology modeling, threading, fragment assembly, and molecular dynamics have been in use for decades. What changed is accuracy and access. Deep learning models that treat a fold as a pattern-recognition problem arrived with force at the CASP14 assessment in 2020, when AlphaFold2 produced predictions close to experimental accuracy for a large share of targets. In the months that followed, laboratories gained a credible structural hypothesis before purifying the protein, before crystallizing it, and sometimes before the gene synthesis order has even been signed off. That shift reverses the usual order of operations, and its full effect is still being sorted out.

From sequence to structure: how prediction models work

Proteins are built from chains of amino acids, and most fold into a preferred three-dimensional shape that determines how they bind other molecules, catalyze reactions, move ions, and respond to signaling partners. A model can't yet solve the physics of folding for every protein, so the best tools take a statistical shortcut. They scan sequences from thousands of related organisms and look for positions that tend to change together. This co-evolution signal indicates residues that sit close in three dimensions even when they are far apart in sequence. The multiple sequence alignment becomes the input, and the network predicts a set of coordinates.

This is why AlphaFold2 and its relatives improved so quickly: the growth of protein sequence databases gave them more evolutionary data. For proteins with deep sequence alignments, the prediction can be sharp. For orphan proteins or synthetic sequences with few relatives, the model has less evidence and the confidence scores drop. The output is therefore not a single answer. It is a hypothesis with a confidence label attached to every region.

The tools now showing up in lab protocols

AlphaFold2 remains the default for many protein-only tasks because its underlying method is described clearly and its predictions are easy to obtain. The AlphaFold2 method paper in Nature laid out the architecture, and the AlphaFold Protein Structure Database now hosts more than 200 million predicted structures from UniProt and other sources. RoseTTAFold, developed at the University of Washington, uses a three-track network that predicts coordinates with lower computational cost and is available through the Robetta server. ESMFold from Meta AI predicts structures quickly from single sequences rather than alignments, which suits metagenomic screening; its approach is described in Science. AlphaFold3 extends the task to protein-DNA, protein-RNA, protein-small-molecule complexes, and post-translationally modified protein forms.

ModelWhat it predictsTypical lab use
AlphaFold2Single-chain protein structuresModel building into cryo-EM maps, construct design, domain boundaries
RoseTTAFoldProtein structures and complexesMonomer and small protein assemblies when compute is limited
ESMFoldProtein structures from single sequencesFast screens of environmental or synthetic proteins
AlphaFold3Proteins with DNA, RNA, ions and ligandsGenerating binding-site and interaction hypotheses

No single tool is best for every problem. A lab studying a conserved enzyme may need only AlphaFold2. A metagenomics group handling millions of fragments may prefer ESMFold for speed. Bench scientists are beginning to choose models the way they choose expression systems: by the question and the budget.

What changes when a structure arrives before the experiment

The most immediate change is in construct design. Crystallographers and cryo-EM researchers have traditionally removed flexible loops or terminal regions by trial and error because unstructured regions prevent lattice contacts or complicate particle alignment. A predicted model with per-residue confidence can identify those regions before a single milligram of protein is produced. Membrane protein labs use predictions to choose truncations that may express better in detergent or nanodiscs. The time saved is not theoretical; it is measured in weeks of failed expression tests.

Predicted models also speed up experimental structure determination. In molecular replacement, a predicted model can serve as a search model when no experimental homolog is available. In cryo-EM, docking a prediction into a modest-resolution map helps trace the polypeptide chain more quickly than building from scratch. Journals and databases have adapted in parallel: Protein Data Bank depositors are now expected to state whether AlphaFold or similar models were used in model building. This is a practical acknowledgment that the two methods are increasingly intertwined.

Reading the confidence scores before trusting the model

The colored confidence plots are the first screen. pLDDT runs from 0 to 100; regions above 90 are generally reliable at the backbone level, while regions below 50 should be treated as unstructured or uncertain. For complexes, the predicted aligned error, or PAE, indicates whether the relative position of two domains is backed by evidence. A low PAE between two regions means the model is confident about their relative orientation; a high PAE means the shared position is a guess.

Researchers should resist the instinct to treat a global confidence score as permission to stop checking. A model can be right about a folded core and wrong about an active-site loop that matters for catalysis. The practical move is to read the per-residue scores first, then inspect the biology: does the predicted pocket line up with known functional residues? Do disulfide bonds form between cysteine residues that should be oxidized? Does the N-linked glycosylation site point outward?

  • Check pLDDT and PAE before using any region in a functional claim.
  • Compare the prediction with experimental data from small-angle X-ray scattering, circular dichroism, cryo-EM maps, or hydrogen-deuterium exchange when available.
  • Confirm critical residues by mutagenesis or binding assays before drawing conclusions.
  • Treat ligand docking positions from predictive models as hypotheses, not measured affinities or binding modes.

AlphaFold3 and the interaction problem

Proteins rarely act alone. The next practical test for molecular structure prediction is not another protein fold but a complex: a transcription factor on DNA, an antibody on an antigen, a kinase with a small-molecule inhibitor, or a viral spike protein with its receptor. AlphaFold3, described in Nature in 2024, extended prediction to these multi-component systems using a diffusion-based approach. It can produce plausible models of protein-ligand, protein-nucleic acid, protein-protein, and antibody-antigen assemblies without a template.

The catch is that plausible is not the same as physical. AlphaFold3 is not a thermodynamic model; it does not measure binding affinity, and it can hallucinate ligand positions in pockets that are only partially formed. The practical use inside a lab is to generate ideas, narrow a screen, or guide medicinal chemistry. The experimental confirmation must still happen, and it usually takes longer than the prediction.

The limits a bench scientist notices

Prediction tools are less reliable for intrinsically disordered regions, flexible loops, membrane protein topology in a lipid-bilayer environment, and proteins that require cofactors or post-translational modifications to fold. They don't model dynamics well: a static structure cannot show conformational changes, allostery, the effect of a mutation on stability, or binding kinetics. A predicted structure of an intrinsically disordered protein often appears as an extended chain with low pLDDT. That isn't a model failure. It is a signal that the protein does not have a stable fold on its own.

Single-point mutations are a related challenge. Most models do not consistently rank which variant will destabilize a protein, and they are not trained to predict temperature sensitivity or aggregation. When a lab asks which of 20 substitutions will break function, the model often answers with confidence intervals too wide to use. Experimental screening remains the backstop.

Validation remains laboratory work

Predictions are now fast enough that the bottleneck has moved to checking them. A structure prediction for a new protein can be generated in an afternoon, but establishing that the model is meaningful may take months: purifying the protein, measuring its oligomeric state, probing a binding site with mutants, or collecting a low-resolution cryo-EM map. The gap between time-to-prediction and time-to-validation is one reason that publications of purely predictive structures are treated with caution by reviewers.

In practice, a lab will combine methods. A prediction can justify a construct. A cryo-EM map can confirm the overall fold. A hydrogen-deuterium exchange experiment can test which regions are flexible. A mutagenesis panel can tie a residue to activity. The model enters this workflow as one piece of evidence, not a substitute for the others.

The question labs will face in year two is straightforward and awkward: did we stop doing the slower experiment because the model was good, or because the model was easy? The answer will vary by protein, by lab and by budget. Better tools can shorten the distance between a sequence and a structure. They can't remove the obligation to test the hypothesis.

a bathroom with a lot of writing on the wall

Photo by Stephan HK on Unsplash

Retrato de Dr. Elena Ramirez
Sobre el autor

Dr. Elena RamirezVer autor

Academic Jobs In House Author

Discusión

por lo menos:

Sé el primero en comentar este artículo!

tú

Se le pedirá que se conecte antes de publicar su comentario.

Nuevo0 comments

¡Únete a la conversación!

¡Añade sus comentarios ahora!

Tenga su palabra

Nivel de compromiso

Browse por Facultad

Browse por tema

Frequently Asked Questions

🧬¿Qué es la predicción de estructura molecular?

La predicción de la estructura molecular es el uso de modelos computacionales para estimar las coordenadas tridimensionales de los átomos en una molécula sin resolver previamente la estructura experimentalmente. Para las proteínas, la entrada suele ser una secuencia de aminoácidos y la salida incluye una forma predicha más puntuaciones de confianza que señalan regiones fiables e inciertas.

🔬¿Cómo predice AlphaFold las estructuras de proteínas?

AlphaFold utiliza aprendizaje profundo e información evolutiva de alineaciones de secuencias múltiples. Busca posiciones de aminoácidos que cambian juntas en proteínas relacionadas, lo que indica residuos que se encuentran cerca en tres dimensiones. El modelo predice entonces un conjunto de coordenadas e informa puntuaciones de confianza para cada región.

📏¿Cuán precisos son los modelos actuales de predicción de estructura molecular?

La precisión depende de la proteína y la cantidad de datos evolutivos. Para muchas proteínas bien conservadas, AlphaFold2 produce predicciones de la columna vertebral cercanas a la precisión experimental. Para proteínas huérfanas, regiones intrínsecamente desordenadas o complejos raros, las puntuaciones de confianza disminuyen y el modelo debe tratarse como una hipótesis en lugar de un resultado.

🧾¿Qué significa pLDDT?

pLDDT significa prueba de diferencia de distancia local predicha. Es una puntuación de confianza por residuo de 0 a 100. Las regiones por encima de 90 suelen ser fiables a nivel de espina dorsal, mientras que las regiones por debajo de 50 deben considerarse no estructuradas o inciertas.

💊¿Pueden los modelos predecir cómo se une una proteína a un fármaco?

AlphaFold3 puede generar modelos de complejos proteína-ligando plausibles, pero no mide la afinidad de unión ni la termodinámica. Una posición de ligando predicha es un punto de partida para experimentos, no un modo de unión medido. Los químicos farmacéuticos aún confirman la unión con ensayos y métodos estructurales.

🧪¿Reemplazan las estructuras predichas la cristalografía de rayos X o la criomicroscopía electrónica?

No. Las predicciones aceleran el diseño de constructos, la creación de modelos y la generación de hipótesis, pero no reemplazan la validación experimental. La cristalografía de rayos X, la criomicroscopía electrónica y la RMN siguen proporcionando restricciones medidas sobre la dinámica, la unión, las modificaciones postraduccionales y la colocación de cadenas laterales.

🧩¿Cuál es la diferencia entre AlphaFold2 y AlphaFold3?

AlphaFold2 predice estructuras de proteínas de cadena única a partir de alineaciones de secuencias. AlphaFold3 amplía la predicción a complejos que contienen proteínas, ADN, ARN, iones y moléculas pequeñas mediante un enfoque basado en difusión. AlphaFold3 es más amplio en alcance, pero requiere una interpretación cuidadosa para modelos de ligandos y multicomponentes.

🖥️¿Qué modelo debe utilizar un laboratorio?

Depende de la pregunta. AlphaFold2 es un valor predeterminado sólido para estructuras de proteínas únicamente. RoseTTAFold es útil cuando el cómputo es limitado o están involucrados complejos. ESMFold es rápido para la selección de muchas secuencias. AlphaFold3 es adecuado para generar hipótesis sobre interacciones con ADN, ARN o moléculas pequeñas.

⚠️¿Cuáles son las principales limitaciones de los modelos de predicción de estructura de proteínas?

Los modelos son menos fiables para regiones intrínsecamente desordenadas, bucles flexibles, topología de proteínas de membrana dentro de una bicapa lipídica, cambios conformacionales, allostery y el efecto de mutaciones puntuales en la estabilidad o la agregación. Producen modelos estáticos y no miden la afinidad de unión.

✅¿Cómo puede un laboratorio validar una estructura predicha antes de su publicación?

Un laboratorio debe comprobar las puntuaciones pLDDT y el error de alineación predicho, comparar el modelo con datos experimentales de SAXS, dicroísmo circular, criomicroscopía electrónica o intercambio de hidrógeno-deuterio, y confirmar los residuos funcionales mediante mutagénesis o ensayos de unión antes de sacar conclusiones.

🗂️¿Se aceptan estructuras predichas en el Banco de Datos de Proteínas?

La Protein Data Bank almacena estructuras determinadas experimentalmente, mientras que la Base de datos de estructuras de proteínas AlphaFold aloja modelos predichos. Las estructuras predichas pueden apoyar la construcción de modelos experimentales, pero se espera que los depositantes indiquen si se utilizaron modelos computacionales y proporcionen la evidencia experimental subyacente.

👩‍🔬¿Necesito ser un experto en aprendizaje automático para ejecutar estos modelos?

No. Muchos modelos están disponibles a través de servidores web como la base de datos AlphaFold, Robetta y cuadernos ColabFold. Los investigadores con experiencia básica en bioinformática pueden ejecutar predicciones, aunque interpretar las puntuaciones de confianza y limitaciones aún requiere juicio de biología estructural.