The Department of Biomedical Data Science (DBDS) at Stanford University held its annual retreat from September 15 to 17 at the Computing and Data (CODA) Science building and E.D. Stone-Edwards Building. Attendees welcomed new and returning students, explored research across the department and celebrated the DBDS community. Keynote speaker Ida Sim, MD’90, PhD’97, BS’85, professor of medicine at the University of California, San Francisco, engaged in a fireside chat with Sylvia Plevritis, chair of biomedical data science and the William M. Hume Professor in the School of Medicine. Attendees also listened to faculty talks, a state of the field panel, faculty debates, and student poster presentations.



State of DBDS
Plevritis opened the retreat with a presentation about the department. Fifty-four Stanford faculty members are affiliated with DBDS, including 19 who are primary faculty. The department also is home to 55 doctoral students, and more than 70 master’s degree students. Plevritis discussed the vision of the department. “We’re transforming what’s happening in AI and biotech into a knowledge revolution of the molecular basis of human health and disease, to produce clinically actionable insights and discoveries that advance precision health and medicine,” Plevritis said. She also unveiled the department’s annual report, which explores ground-breaking research, academic achievements and community highlights from the past year.
We’re transforming what’s happening in AI and biotech into a knowledge revolution of the molecular basis of human health and disease, to produce clinically actionable insights and discoveries that advance precision health and medicine. –Sylvia Plevritis


Fireside Chat with Alumni Keynote Speaker Ida Sim

Sim is a global leader in the technology and policy of large-scale health data sharing. As she shared in conversation with Plevritis, her work in this field began with her doctoral thesis, which proposed representing clinical trials as computational, knowledge artifacts. Two decades later, she co-founded Vivli, now the world’s largest clinical trial data-sharing platform. She also was the project coordinator of the World Health Organization’s International Clinical Trials Registry Platform, where she led the establishment of mandatory global clinical trial registration.
Additionally, Sim discussed her leadership at the JupyterHealth project, which securely connects wearable, clinical and patient-generated data into computational and AI-enabled health care systems. Understanding that 75 percent of Americans live with a chronic condition and 60 percent live with two or more, Sim and her colleagues saw a need to bring together patients’ disparate health data under one health system. Speaking directly to students in the audience, Sim advised that doing this work was about following her North Star. “What is your North Star?” she said. “For me, it’s impacting real people in the real world. They need you more than ever.”
Continuing with a human-centered message, Sim also pushed back on the “human in the loop” framework in the context of AI in healthcare and clinical research. “We say a lot about the human in the loop, and I really object to that, because that assumes there is a loop that’s already defined. No, the entire loop is human, what we do as humans, and our question is, ‘Where does the AI fit in?’” Sim said.
At the end of her discussion, Sim counseled students to continuously immerse themselves in real-world settings, such as shadowing clinicians and working in community health, and to learn to ask good questions by listening to those asked by colleagues. “How can you learn to push their question to the next question? Get into the habit of that,” Sim said.
What is your North Star? For me, it’s impacting real people in the real world. They need you more than ever.–Ida Sim


State of the Field Panel


The State of the Field panel focused on three questions: where the field of biomedical data science stood five years ago, where it is now, and where it is headed.
Chiara Sabatti, professor of biomedical data science and of statistics, said that what it means to do science—particularly when multi-agent systems can formulate endless hypotheses—is changing. She cautioned, though, that it is difficult to make predictions for a field that is advancing so rapidly because of AI. “It feels like we are in some primordial chaos, and we’ll see what happens next. It’s disorienting because we cannot continue to do things as we have done, but it’s also very exciting because a lot of opportunities are opening up for us,” Sabatti said.
The conversation expanded into a broader debate about AI’s trajectory in biomedical science, including whether AI development should slow down. The panelists largely agreed that slowing down is not realistic; rather, AI needs to be developed with better guardrails, evaluation standards and regulation. Tina Hernandez-Boussard, associate dean of research and professor of medicine and of biomedical data science, raised the challenge of tracking and stopping errors that might start with one agent but propagate through chains of interacting agents in a clinical setting. “It’s no longer one person who is accountable,” Hernandez-Boussard said. “It’s distributed accountability across an entire team.”
Sabatti and Hernandez-Boussard were joined on the panel by their colleagues, Aaron Newman and James Zou, associate professors of biomedical data science. The panelists also emphasized the importance of biomedical domain expertise of humans in the age of AI and the value of academia versus industry in having access to the data—from basic science to clinical application—that is necessary for training models.
In their closing thoughts, the panelists agreed that teaching students how to evaluate the quality of scientific questions—particularly those posed by AI—takes on increased importance as scientists continue to collaborate with AI agents. Furthermore, because of the nature of this collaboration, the panelists suggested that the norms around credit attribution, authorship and peer review will need to evolve.


Faculty Debates



The retreat also featured three faculty debates on how AI is reshaping biomedical data science.
Rob Tibshirani, professor of biomedical data science, and Curtis Langlotz, professor of radiology, debated AI’s place in education. Tibshirani argued that in-person learning remains essential to foster students’ critical thinking and problem-solving capabilities. Langlotz compared AI to earlier technologies, like the internet, that are now indispensable to education. “In a few years, we’re not going to think about not using AI,” he said.
Nigam Shah, professor of medicine, and Roxana Daneshjou, assistant professor of biomedical data science, debated how quickly clinical AI should be deployed. Shah argued that excessive regulatory caution is holding back care improvements. Daneshjou countered that AI is often oversold as a cure for a fragmented health system that would benefit just as much from investment in public health. Both agreed that debates over AI replacing physicians miss the point. “I think we confuse tasks with jobs,” Shah said.
Alexander Ioannidis, assistant professor of genetics and of biomedical data science, and John Witte, professor of epidemiology and population health and of biomedical data science, discussed AI’s role in genomics, including the somewhat surprising finding that polygenic risk score models are not substantively improved by incorporating gene-gene interactions.
On the topic of people using AI to screen for genetic traits to create “designer babies,” Ioannidis and Witte agreed that academia has a role to play in critiquing the accuracy of claims by companies offering these type of services and in addressing the ethical implications of various types of screening.



Faculty Talks


Serena Yeung-Levy, assistant professor of biomedical data science, presented her lab’s work teaching AI models to reason over biological and clinical images, from microscopy to tissue scans. Her lab built MicroVQA, a rigorously vetted benchmark of expert-level microscopy questions now used across academia and industry. “There was a sense that frontier models can do a lot,” Yeung-Levy said, “and yet experimentalists…would often have the sense of, they’re saying things that are not incorrect, but maybe are still not useful enough.” Her lab has released a 24-million-image, open-access dataset drawn from PubMed Central to help train future biomedical vision-language models.
Jason Fries, assistant professor of biomedical data science, discussed his lab’s work of building “patient world models,” in the context of oncology, that simulate how a patient’s health might unfold over time. His lab’s generative model, trained on 4.3 million patient records, generates multiple plausible future trajectories for a patient. He argued that closing the gap to clinical reality is a systems challenge as much as a modeling one. “It is how we connect models to users to data to researchers,” Fries said.
Twelve faculty members also delivered four-minute lightning talks: Daneshjou, Ioannidis, Langlotz, Plevritis, Sabatti, Witte, Emily Alsentzer, assistant professor of biomedical data science; Mark Musen, professor of medicine; Julia Palacios, associate professor of statistics; Manuel Rivas, assistant professor of biomedical data science; Dennis Wall, professor of pediatrics; and Leila Wehbe, associate professor of biomedical data science.






Student Posters
More than fifty students and postdoctoral fellows registered to present research posters. Those presented by Marie Amale Huynh, Philip Adamson, Rebecca Hurwitz and Dehua Bi represent a wide range of complex research questions and problems that young researchers in the department are tackling.
Huynh presented NEST, believed to be the first compositional benchmark for nonverbal social reasoning in child-centric videos, aimed at supporting early child development research. Adamson, a machine learning engineer, discussed VISTA, an ARPA-H-funded project integrating longitudinal health records, imaging, pathology and outcomes data for hundreds of thousands of Stanford cancer patients. Hurwitz shared a model that demonstrates how combined environmental exposures, such as air pollution and the built environment, affect conditions like asthma over a lifetime. Bi, a postdoctoral scholar, presented a method to make clinical trials more reliable when patients can’t be split evenly, using computer-simulated “digital twins” of patients.







Llamas, Scavenger Hunts, and a Road Trip to Pacifica
Having fun is also an important goal of DBDS Retreats. This year, attendees were entertained by llamas, a campus-wide scavenger hunt, and tacos and drinks at the famous beach-side Pacifica Taco Bell.







