Researchers at Stanford Medicine have created two new artificial intelligence models that simulate biological cells. The first model, called "universal cell embedding," laid the groundwork for the second, more advanced model named TranscriptFormer. This second model was trained using data from 112 million cells across 12 different species, ranging from simple single-celled organisms like yeast to humans. These AI models make it easier to compare cells across species and could lead to new insights into diseases and potential cell-based treatments. "We're trying to open up a whole new way to think about cell biology," said Stephen Quake, Ph.D., a professor of bioengineering and co-lead author of two scientific papers on the work, published in Nature and Science. The genome contains the instructions for life, but many biologists now focus on gene expression — which genes a cell is actively using. For example, beta cells in the pancreas express genes for insulin and related functions, while white blood cells called B cells express genes involved in fighting disease. Different cell types, such as skin cells, express different sets of genes, which helps scientists distinguish them. Similarly, healthy and diseased cells, as well as cells from different species, have distinct gene expression patterns. In recent years, scientists have built large databases, or "atlases," that map these patterns across many species. However, the sheer volume of data — involving tens of thousands of genes from hundreds of millions of cells — is overwhelming for human analysis. To manage this complexity, researchers developed TranscriptFormer, which was trained on data from 12 species, including humans, yeast, mice, rabbits, chickens, zebrafish, fruit flies, frogs, malaria-causing parasites, sea urchins, sponges, and the roundworm Caenorhabditis elegans. Like large language models such as ChatGPT, which learn by predicting missing words in text, TranscriptFormer learned by analyzing patterns in gene expression data. This training allowed the model to create a "universal space" — a mathematical framework where cells from all species can be compared based on their gene expression profiles. This model has several practical applications. For example, it can help scientists explore evolutionary relationships by identifying similarities between cells in different species. Researchers studying sponges, for instance, used the model to determine which sponge cell is most similar to neurons in other animals. The model also helped clarify the function of certain sponge cells, revealing that some previously thought to be related to neurons might instead be involved in digestion. Additionally, TranscriptFormer can identify cell types in new species and distinguish between healthy and diseased cells. In the future, it could aid in designing new cell-based treatments by predicting the properties of cells that currently do not exist. "We hope it's going to be a very powerful tool for discovery," Quake said. The research was supported by the Chan Zuckerberg Initiative and BioHub, with contributions from researchers at Stanford and other institutions.