
Claudio Spiess
PhD Candidate, Computer Science
Bio
Hey there! I’m Claudio 👋🏼 I’m a PhD candidate in Computer Science at the University of California, Davis, affiliated with the DECAL Lab. My research interests are in the intersection of Natural Language Processing and Software Engineering. I study how machine learning techniques, mostly from the NLP world, can understand source code, and to a greater extent, help software engineering. In particular, most of my work is focused on applying LLMs (Large Language Models) to source code. I’m fortunate to be advised by Prof. Prem Devanbu, and to have collaborated with exceptional colleagues around the world. My work has been published at flagship venues such as ICSE and FSE, and I have served as a reviewer for premier journals such as TOSEM.
From model “understanding”, we can do useful things like program generation from natural language, automated bug fixing, automated documentation, anomaly detection (bugs!), reverse engineering, naming, among many others. Not only do I seek to build systems, but to investigate metaphysical questions: do LLMs understand code? What do they learn? What biases and problems do these approaches have? And most importantly, how to fix them? On the practical side, I’m interested in how cognitive load can be alleviated while writing software by smart tools for programmers. Between 2023 and 2025, my main focus was on the calibration of LLMs for code, or rather the lack thereof. In 2025, I also worked on prompt programming languages and automated prompt optimization for LLM agents. Recently, I have been working on how LLMs understand and reason about code.
Previously, I helped build a data driven lending platform at Dutch FinTech startup Floryn as a full stack software engineer, making machine learning work for loans. Most processes had some form of machine learning algorithms backing them, so I built interesting interpretation and explanation tools for non-technicals.
I received my bachelor degree in Computer Science & Engineering from the Free University of Bolzano. I wrote a research thesis, concentrating on NLP for software engineering, under the supervision of Dr. Romain Robbes and Dr. Andrea Janes. During this time, I was affiliated with the Software and Systems Engineering (SwSE) research group.
When I’m not hacking around on code or models, I like to travel the world with a backpack (44 countries/territories and counting), scuba dive (150 dives and counting), and hike volcanoes. I also speak six languages: English, German, French, Dutch, Italian, and some Spanish.
Publications

How Robustly Do LLMs Understand Execution Semantics?

Does In-IDE Calibration of Large Language Models work at Scale?

Model See, Model Do? Exposure-Aware Evaluation of Bug-vs-Fix Preference in Code LLMs

On LLMs’ Internal Representation of Code Correctness

Calibration and Correctness of Language Models for Code

AutoPDL: Automatic Prompt Optimization for LLM Agents

STraceBERT: Source Code Retrieval using Semantic Application Traces
Projects

PDL: A Declarative Prompt Programming Language

Method Name Suggestions: An Open Vocabulary Approach

Universal Transformer: Towards Learned Positional Encodings

Impact of War on Food Security in the Middle East & East Africa

Simulating crop yields in El Oro, Ecuador

SPM: Social Package Manager
pandas-dp: Differential Privacy in pandas

CUDA Docker Stack

Kickstarter Project Analysis
Subtitle Keyword Extractor
PhotoStack
BBQ-Planner
DNA analysis

Jtrak to Macdive
ToodleMoodle


Address Book

ApiCollider

Natürliche Rechenmaschine
Network Dropbox
Cite How Robustly Do LLMs Understand Execution Semantics?
@inproceedings{spiess2026execution,
title = {{How Robustly Do LLMs Understand Execution Semantics?}},
author = {Spiess, Claudio and Devanbu, Prem and Barr, Earl T.},
year = {2026},
booktitle = {Proceedings of the 3rd ACM International Conference on AI-Powered Software},
pages = {288--298},
publisher = {ACM},
url = {https://doi.org/10.1145/3805760.3814919},
doi = {10.1145/3805760.3814919},
}Cite Does In-IDE Calibration of Large Language Models work at Scale?
@inproceedings{koohestani2026calibration,
title = {{Does In-IDE Calibration of Large Language Models work at Scale?}},
author = {Koohestani, Roham and Sergeyuk, Agnia and Gros, David and Spiess, Claudio and Titov, Sergey and Devanbu, Premkumar and Izadi, Maliheh},
year = {2026},
booktitle = {Proceedings of the 34th ACM International Conference on the Foundations of Software Engineering},
pages = {609--619},
publisher = {ACM},
url = {https://doi.org/10.1145/3803437.3805234},
doi = {10.1145/3803437.3805234},
}Cite Model See, Model Do? Exposure-Aware Evaluation of Bug-vs-Fix Preference in Code LLMs
@inproceedings{alkaswan2026exposure,
title = {{Model See, Model Do? Exposure-Aware Evaluation of Bug-vs-Fix Preference in Code LLMs}},
author = {Al-Kaswan, Ali and Spiess, Claudio and Devanbu, Prem and van Deursen, Arie and Izadi, Maliheh},
year = {2026},
booktitle = {Proceedings of the 23rd International Conference on Mining Software Repositories},
pages = {86--97},
publisher = {ACM},
url = {https://doi.org/10.1145/3793302.3793341},
doi = {10.1145/3793302.3793341},
}Cite On LLMs’ Internal Representation of Code Correctness
@inproceedings{ribeiro2026correctness,
title = {{On LLMs’ Internal Representation of Code Correctness}},
author = {Ribeiro, Francisco and Spiess, Claudio and Devanbu, Premkumar and Nadi, Sarah},
year = {2026},
booktitle = {Proceedings of the 2026 IEEE/ACM 48th International Conference on Software Engineering},
pages = {1160--1172},
publisher = {ACM},
url = {https://doi.org/10.1145/3744916.3787846},
doi = {10.1145/3744916.3787846},
}Cite Calibration and Correctness of Language Models for Code
@inproceedings{spiess2025calibration,
title = {{Calibration and Correctness of Language Models for Code}},
author = {Spiess, Claudio and Gros, David and Pai, Kunal Suresh and Pradel, Michael and Rabin, Md Rafiqul Islam and Alipour, Amin and Jha, Susmit and Devanbu, Prem and Ahmed, Toufique},
year = {2025},
booktitle = {2025 IEEE/ACM 47th International Conference on Software Engineering (ICSE)},
pages = {540--552},
publisher = {IEEE},
url = {https://doi.org/10.1109/ICSE55347.2025.00040},
doi = {10.1109/ICSE55347.2025.00040},
}Cite AutoPDL: Automatic Prompt Optimization for LLM Agents
@inproceedings{pmlr-v293-spiess25a,
title = {{AutoPDL: Automatic Prompt Optimization for LLM Agents}},
author = {Spiess, Claudio and Vaziri, Mandana and Mandel, Louis and Hirzel, Martin},
year = {2025},
booktitle = {Proceedings of the Fourth International Conference on Automated Machine Learning},
pages = {13/1--20},
volume = {293},
series = {Proceedings of Machine Learning Research},
publisher = {PMLR},
editor = {Akoglu, Leman and Doerr, Carola and van Rijn, Jan N. and Garnett, Roman and Gardner, Jacob R.},
month = {September},
pdf = {https://raw.githubusercontent.com/mlresearch/v293/main/assets/spiess25a/spiess25a.pdf},
url = {https://proceedings.mlr.press/v293/spiess25a.html},
}Cite STraceBERT: Source Code Retrieval using Semantic Application Traces
@inproceedings{spiess2023stracebert,
title = {{STraceBERT: Source Code Retrieval using Semantic Application Traces}},
author = {Spiess, Claudio},
year = {2023},
booktitle = {Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering},
pages = {2207--2209},
publisher = {ACM},
url = {https://doi.org/10.1145/3611643.3617852},
doi = {10.1145/3611643.3617852},
}Cite PDL: A Declarative Prompt Programming Language
@misc{vaziri2024pdldeclarativepromptprogramming,
title={PDL: A Declarative Prompt Programming Language},
author={Mandana Vaziri and Louis Mandel and Claudio Spiess and Martin Hirzel},
year={2024},
eprint={2410.19135},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2410.19135},
}
