[lamarr-nlp] Guest Talk by Jan Philip Wahle from University of Goettingen | Can We Trust What Models Think and Say?
As part of the Lamarr NLP Colloquium, we have the pleasure to host Jan Philip Wahle from University of Goettingen. Jan will give a talk on the trustworthiness of language models.
Title: Can We Trust What Models Think and Say?
Abstract:
Can we trust a model simply because its outputs look correct, harmless, or well-reasoned? As language models become more capable, this question becomes increasingly difficult to answer. Undesirable behavior may remain hidden in seemingly benign outputs, while verbalized reasoning may not faithfully reflect the computations that drive a model’s decisions. In this talk, I will examine what it means to trust model behavior and how that trust can be measured. Using covert information leakage and reasoning faithfulness as two case studies, I will discuss where behavioral evaluation breaks down, what chain-of-thought can and cannot reveal, and how internal representations can help us distinguish safe-looking behavior from genuinely trustworthy model behavior.
Date: Wednesday, September 9, 2026
Time: 11:00 - 12:00 pm (CET).
Location: Friedrich-Hirzebruch-Allee 6, 53115 Bonn, Germany
Room: 2.122 + Zoom
Zoom: https://uni-bonn.zoom-x.de/j/63819604806?pwd=64PSGa9HyTym9j1bjy6jhcJF3eHebi.1