Structure prediction

Preview

NovaFold-Lite-Preview

Every heavy atom of an assembly, from sequence — proteins, nucleic acids, and the molecules they bind, folded together.

About this release

NovaFold-Lite-Preview is the first model in the NovaFold family, offering an early look at our next generation of protein structure prediction models. Designed as a lightweight and efficient model, it enables fast inference while establishing the foundation for the larger NovaFold models that follow.

We are actively scaling the NovaFold family across model capacity and compute. Larger models are currently in training, alongside expanded benchmarking and calibration studies.

NovaFold-Lite-Preview represents the first step in this model family. Additional models, evaluations, and capabilities will be released progressively.

Training data cutoff
30 Sep 2021
Molecule classes
4 jointly modelled
Overview

An all-atom structure predictor.

Given the sequences of a biomolecular assembly, novafold predicts the position of every heavy atom in it, together with calibrated confidence for each residue and each pair of residues.

It pairs a recycled pairformer trunk with a diffusion process over atoms: the trunk reads sequence, evolutionary alignments, and optional structural templates to build a representation of the whole assembly, and the diffusion model generates coordinates conditioned on it.

Molecule classes

Everything modelled together.

Classes aren't predicted separately and assembled afterward. A protein–DNA complex with a bound cofactor is one prediction, and the interfaces between its parts are modelled as part of it.

Protein

Any number of chains, with post-translational and non-standard residues named per position.

DNA

Single- and double-stranded, modelled in the same pass as its protein partners.

RNA

Structured RNAs and protein–RNA complexes, from sequence.

Ligand

Small molecules, cofactors, metals, and ions — by chemical component code or SMILES.

What you can ask for

Constraints the model was trained on.

Constraints enter as conditioning, not as a filter applied afterward.

Cyclic peptides & macrocycles

Declare a chain cyclic and the model treats it as a closed ring — the two ends read as neighbours, not distant termini.

Covalent modifications

Name two atoms and novafold bonds them: a covalent inhibitor to its catalytic residue, a glycan to its sequon, a crosslink between chains.

Pocket & contact constraints

Point a binder at a known site by naming the residues it should contact, and the distance you mean it to hold.

What comes back

Structure, plus how sure it is.

Coordinates

All heavy atoms, five independent samples per target, ranked.

pLDDT

Per-residue confidence in the local structure.

PAE

Expected error between any two residues once aligned.

PDE

Expected error in the distance between two residues.

pTM / ipTM

Whole-complex and per-interface reliability, per chain pair.

Distogram

The full distance distribution, not just the chosen structure.