A new AI model called CLSS (Contrastive Learning Sequence-Structure) has achieved what was once thought impossible: it fuses the two fundamental languages of proteins—amino acid sequence and three-dimensional structure—into a single, unified representation. Published today in Proceedings of the National Academy of Sciences, this breakthrough promises to give biologists a bird’s-eye view of the entire protein universe, revealing how millions of protein families are related across billions of years of evolution.

What Happened

An international team led by Professor Rachel Kolodny and PhD candidate Guy Yanai at the University of Haifa, together with Professor Nir Ben-Tal and graduate student Gabriel Axel at Tel Aviv University, and Specially Appointed Associate Professor Liam M. Longo at the Earth-Life Science Institute (ELSI) at Institute of Science Tokyo, created CLSS. The name stands for Contrastive Learning Sequence-Structure—a deep learning approach that was trained on both sequence and structural data simultaneously.

The key innovation is the model’s ability to align protein sequences with their corresponding 3D folds in a common embedding space. This allows researchers to traverse the protein universe continuously, finding evolutionary transitions and distant relationships that were previously hidden. CLSS can be used to predict protein function, infer evolutionary history, and even design novel proteins with desired properties. Kolodny spent five months as a visiting researcher at ELSI developing methods to analyze the model’s outputs.

The paper draws on the wealth of structural data accumulated by cryo-EM and alphafold-era efforts, turning that static library into a dynamic map. By treating sequence and structure as two views of the same object and learning to contrast them, CLSS effectively builds a coordinate system for the entire protein universe.

Read the full announcement →

My Take

This is a landmark. While AlphaFold solved structure prediction, it didn’t give us a unified language for all proteins. CLSS does. For bioinformatics and drug discovery, this is like having a Google Maps for the protein world—you can now ask “what’s the nearest functional relative to this unknown protein?” and get an answer grounded in both sequence and shape.

What excites me most is the evolutionary angle. The protein universe has always been studied in fragments—sequence trees and structure classifications rarely talked to each other. CLSS bridges that gap. For developers, this opens up new possibilities in protein engineering, AI-guided drug design, and even synthetic biology. The model’s architecture (contrastive learning on multimodal data) is also a template that could be applied to other biological languages, like RNA or metabolites.

What to Watch

  • Functional annotation: CLSS can likely assign functions to “dark” proteins much more accurately than sequence-only methods. Expect a wave of new annotations on UniProt.
  • Protein design: By navigating the unified embedding space, researchers can generate sequences that fold into desired structures—a direct route to designer enzymes and therapeutics.
  • Cross-domain transfer: The contrastive learning approach used here—pairing two modalities into one embedding—may soon be adopted for other biology problems, such as linking genotypes to phenotypes or metabolomics to proteomics.
  • Open-source availability: The paper hints at methods being developed at ELSI. If the team releases pretrained models, expect rapid adoption in academic and industrial labs.