Predict a protein's folding rate from its sequence alone
In plain words
Small proteins fold at rates that differ by roughly a factor of a hundred million, and sequence-based formulas reproduce the trend only on proteins they were fitted to. The goal is a sequence-only predictor whose accuracy holds on proteins it has never seen.
Precise statement
For single-domain two-state proteins of 40 to 150 residues in water near $298\,\mathrm{K}$, predict $\ln k_f$ ($k_f = \text{folding rate in } \mathrm{s}^{-1}$) from sequence alone; measured rates span approximately 8 decades, topology measures such as relative contact order correlate with $\ln k_f$ at $r \sim 0.8$ when the native structure is given, and sequence-only fits reach $r \sim 0.8$ in-sample (Ivankov and Finkelstein, 2004). An answer is a method with a stated root-mean-square error, target below 1 unit of $\ln k_f$, on a blind set of proteins with measured rates that are absent from its training data.
What would settle it
A blind test on newly measured folding rates of proteins absent from any training data, scored by root-mean-square error in $\operatorname{ln} k_f$.
Status in the literature
Unverified note
Structure-prediction networks (AlphaFold 2 in 2021 and successors) give native structures, but no blind sequence-to-rate benchmark at this accuracy is known to this survey as of 2026.