Evento
#Seminarios DaSCI

“Mechanistically Interpreting Language Models: From Syntactic Code Completion to Arithmetic Reasoning”

A series of seminars on Artificial Intelligence and Cybersecurity, part of the IAFER Chair project, will host the conferenceMechanistically Interpreting Language Models: From Syntactic Code Completion to Arithmetic Reasoning
The talk will be given by D. Ziyu Yao, Assistant Professor in the Department of Computer Science at George Mason University.

Online Room:  https://oficinavirtual.ugr.es/redes/SOR/SALVEUGR/accesosala.jsp?IDSALA=22999060
Password: 180006

Abstract:
Transformer-based language models (LMs) have shown remarkable progress in tackling complex tasks. However, their rapid advancements come with growing concerns about safety and reliability, stemming from our limited understanding of their inner workings. In this talk, I will share our recent efforts toward building a mechanistic understanding of LMs. First, I will present our study on why LMs struggle with a seemingly simple syntactic code completion task: balanced parentheses completion. We discovered that LMs fail not because of their lack of reliable mechanisms to accomplish the task, but that the faulty mechanisms inside them overshadow the sound ones. Next, I will share our findings about how LMs perform arithmetic on expressions such as “a + b – c”. Surprisingly, we found that LMs do not solve these equations in a human-like compositional manner. Instead, they transfer all information to the last-token position and rely on it to complete the calculation. Finally, I will conclude the talk with a brief discussion of our vision on mechanistically interpreting LMs.

Bio:
Ziyu Yao (https://ziyuyao.org/) is an Assistant Professor in the Department of Computer Science at George Mason University, where she co-leads the George Mason NLP group (https://nlp.cs.gmu.edu/). She works on LLM reasoning, planning, mechanistic interpretability, and human-LLM interaction, and has organized workshops on these topics at COLM, ACL, and NAACL. She recently gave a tutorial about Mechanistic Interpretability of LMs at ICML 2025. Her work has been funded by National Science Foundation, Virginia CCI, and Microsoft, among others. Prior to George Mason, she graduated with a Ph.D. degree in Computer Science and Engineering from the Ohio State University in 2021.

Conference Recording