An international team of scientists, including researchers from the Earth-Life Science Institute (ELSI) at the Institute of Science Tokyo, has created a new type of artificial intelligence model that helps understand the vast world of proteins. This model, called Contrastive Learning Sequence-Structure (CLSS), combines information about the sequence of amino acids that make up a protein and its three-dimensional shape. This approach allows scientists to better understand how proteins are related to each other and how they have evolved over billions of years. The research was led by scientists from the University of Haifa, Tel Aviv University, and ELSI, and the findings were published in the Proceedings of the National Academy of Sciences. Proteins are essential for nearly all functions inside living cells, and understanding their relationships is a key question in evolutionary biology. Scientists have long classified proteins into groups based on their similarities, much like how species are grouped in biology. However, advances in artificial intelligence are now offering new ways to explore these relationships. One such method involves "protein language models," which convert proteins into numerical representations called "embeddings." These embeddings can be thought of as a kind of address or code that groups similar proteins together, helping scientists visualize their relationships in a "protein world map." A challenge in this work is that proteins contain information in both their amino acid sequences and their three-dimensional structures. These two aspects are not always directly related—proteins with different sequences can have similar structures, and similar sequences can lead to very different structures. Most existing models have treated sequence and structure separately, making it difficult to see how they relate. The CLSS model, however, was specifically designed to bring these two aspects together. It uses a technique called contrastive learning, where the model is trained to produce similar embeddings for both the sequence and structure of the same protein, ensuring they appear close to each other on the map. When tested, CLSS outperformed other models in creating a unified map that accurately reflected known protein relationships. It also successfully captured patterns seen in traditional classification systems, even though it wasn’t trained on those systems. The model’s ability to integrate sequence and structure information also made it more effective in classification tasks. Researchers say this approach could help uncover evolutionary patterns that are hard to detect using conventional methods. Additionally, CLSS can use short fragments of protein sequences to create meaningful representations, which is especially useful for studying evolutionary history, as small protein pieces have been reused and rearranged over time. The maps created by CLSS also revealed broader patterns in protein space. When biological properties such as the presence of certain molecules were overlaid onto the maps, distinct groupings emerged—like proteins that bind to metals being spread out more than those that bind to organic molecules. These findings show that such maps can not only classify proteins but also help scientists explore how their structure, function, and evolutionary history are connected. The researchers believe that this kind of unified representation could lead to new advances in protein research, including better database searches, protein engineering, and understanding how the diversity of proteins seen today evolved over billions of years.