Publications

Claudio Spiess,
Prem Devanbu,
Earl T. Barr
How Robustly Do LLMs Understand Execution Semantics?
How Robustly Do LLMs Understand Execution Semantics?
AIware 2026,
2026
We study the robustness of LLM program-output prediction under code transformations and input perturbations, including inputs that raise exceptions. The results expose failures hidden by performance on the original benchmark.

Roham Koohestani,
Agnia Sergeyuk,
David Gros,
Claudio Spiess,
Sergey Titov,
Premkumar Devanbu,
Maliheh Izadi
Does In-IDE Calibration of Large Language Models work at Scale?
Does In-IDE Calibration of Large Language Models work at Scale?
FSE 2026,
2026
Collaboration with JetBrains Research Prior work suggests calibration can improve alignment, but at-scale evidence is limited. In this work, we investigate the feasibility of applying calibration of code models to an in-IDE context. We study two aspects of the problem: (1) the technical method for implementing confidence calibration and improving the reliability of code generation models, and (2) the human-centered design principles for effectively communicating reliability signal to developers.

Ali Al-Kaswan,
Claudio Spiess,
Prem Devanbu,
Arie van Deursen,
Maliheh Izadi
Model See, Model Do? Exposure-Aware Evaluation of Bug-vs-Fix Preference in Code LLMs
Model See, Model Do? Exposure-Aware Evaluation of Bug-vs-Fix Preference in Code LLMs
MSR 2026,
2026
We study how training-data exposure to buggy and fixed code affects code LLM preferences and generations, using membership testing to distinguish memorization from generalization.

Francisco Ribeiro,
Claudio Spiess,
Premkumar Devanbu,
Sarah Nadi
On LLMs’ Internal Representation of Code Correctness
On LLMs’ Internal Representation of Code Correctness
ICSE 2026,
2026
Inspired by findings that LLMs internally encode concepts like truthfulness, this paper explores if LLMs similarly represent code correctness. Specifically, we identify a correctness representation inside LLMs by contrasting the hidden states between pairs of correct and incorrect code for the same programming tasks. By experimenting on four LLMs, we show that exploiting this extracted correctness representation outperforms standard log-likelihood ranking, as well as verbalized model confidence. Furthermore, we explore how this internal correctness signal can be used to select higher-quality code samples, without requiring test execution.

Claudio Spiess,
David Gros,
Kunal Suresh Pai,
Michael Pradel,
Md Rafiqul Islam Rabin,
Amin Alipour,
Susmit Jha,
Prem Devanbu,
Toufique Ahmed
Calibration and Correctness of Language Models for Code
Calibration and Correctness of Language Models for Code
ICSE 2025,
2025
Machine learning models are widely used but can also often be wrong. Users would benefit from a reliable indication of whether a given output from a given model should be trusted, so a rational decision can be made whether to use the output or not. In this work, we studied whether current LLMs for code are well calibrated.

Claudio Spiess,
Mandana Vaziri,
Louis Mandel,
Martin Hirzel
AutoPDL: Automatic Prompt Optimization for LLM Agents
AutoPDL: Automatic Prompt Optimization for LLM Agents
AutoML 2025,
2025
Work done at IBM Research This paper proposes AutoPDL, an automated approach to discover good LLM agent configurations. Our method frames this as a structured AutoML problem over a combinatorial space of agentic and non-agentic prompting patterns and demonstrations, using successive halving to efficiently navigate this space. We introduce a library implementing common prompting patterns using the PDL prompt programming language. AutoPDL solutions are human-readable, editable, and executable PDL programs that use this library.

Claudio Spiess
STraceBERT: Source Code Retrieval using Semantic Application Traces
STraceBERT: Source Code Retrieval using Semantic Application Traces
ESEC/FSE 2023,
2023
Novel approach that utilizes a Java dynamic analysis tool to record calls to core Java libraries, and a BERT-style model on the recorded application traces for effective method source code retrieval from a candidate set. Experiments demonstrate the effectiveness in retrieving the source code compared to existing approaches. Proposed approach offers a promising solution to the problem of code retrieval in software reverse engineering and opens up new avenues for further research in this area.
Cite How Robustly Do LLMs Understand Execution Semantics?
@inproceedings{spiess2026execution,
title = {{How Robustly Do LLMs Understand Execution Semantics?}},
author = {Spiess, Claudio and Devanbu, Prem and Barr, Earl T.},
year = {2026},
booktitle = {Proceedings of the 3rd ACM International Conference on AI-Powered Software},
pages = {288--298},
publisher = {ACM},
url = {https://doi.org/10.1145/3805760.3814919},
doi = {10.1145/3805760.3814919},
}Cite Does In-IDE Calibration of Large Language Models work at Scale?
@inproceedings{koohestani2026calibration,
title = {{Does In-IDE Calibration of Large Language Models work at Scale?}},
author = {Koohestani, Roham and Sergeyuk, Agnia and Gros, David and Spiess, Claudio and Titov, Sergey and Devanbu, Premkumar and Izadi, Maliheh},
year = {2026},
booktitle = {Proceedings of the 34th ACM International Conference on the Foundations of Software Engineering},
pages = {609--619},
publisher = {ACM},
url = {https://doi.org/10.1145/3803437.3805234},
doi = {10.1145/3803437.3805234},
}Cite Model See, Model Do? Exposure-Aware Evaluation of Bug-vs-Fix Preference in Code LLMs
@inproceedings{alkaswan2026exposure,
title = {{Model See, Model Do? Exposure-Aware Evaluation of Bug-vs-Fix Preference in Code LLMs}},
author = {Al-Kaswan, Ali and Spiess, Claudio and Devanbu, Prem and van Deursen, Arie and Izadi, Maliheh},
year = {2026},
booktitle = {Proceedings of the 23rd International Conference on Mining Software Repositories},
pages = {86--97},
publisher = {ACM},
url = {https://doi.org/10.1145/3793302.3793341},
doi = {10.1145/3793302.3793341},
}Cite On LLMs’ Internal Representation of Code Correctness
@inproceedings{ribeiro2026correctness,
title = {{On LLMs’ Internal Representation of Code Correctness}},
author = {Ribeiro, Francisco and Spiess, Claudio and Devanbu, Premkumar and Nadi, Sarah},
year = {2026},
booktitle = {Proceedings of the 2026 IEEE/ACM 48th International Conference on Software Engineering},
pages = {1160--1172},
publisher = {ACM},
url = {https://doi.org/10.1145/3744916.3787846},
doi = {10.1145/3744916.3787846},
}Cite Calibration and Correctness of Language Models for Code
@inproceedings{spiess2025calibration,
title = {{Calibration and Correctness of Language Models for Code}},
author = {Spiess, Claudio and Gros, David and Pai, Kunal Suresh and Pradel, Michael and Rabin, Md Rafiqul Islam and Alipour, Amin and Jha, Susmit and Devanbu, Prem and Ahmed, Toufique},
year = {2025},
booktitle = {2025 IEEE/ACM 47th International Conference on Software Engineering (ICSE)},
pages = {540--552},
publisher = {IEEE},
url = {https://doi.org/10.1109/ICSE55347.2025.00040},
doi = {10.1109/ICSE55347.2025.00040},
}Cite AutoPDL: Automatic Prompt Optimization for LLM Agents
@inproceedings{pmlr-v293-spiess25a,
title = {{AutoPDL: Automatic Prompt Optimization for LLM Agents}},
author = {Spiess, Claudio and Vaziri, Mandana and Mandel, Louis and Hirzel, Martin},
year = {2025},
booktitle = {Proceedings of the Fourth International Conference on Automated Machine Learning},
pages = {13/1--20},
volume = {293},
series = {Proceedings of Machine Learning Research},
publisher = {PMLR},
editor = {Akoglu, Leman and Doerr, Carola and van Rijn, Jan N. and Garnett, Roman and Gardner, Jacob R.},
month = {September},
pdf = {https://raw.githubusercontent.com/mlresearch/v293/main/assets/spiess25a/spiess25a.pdf},
url = {https://proceedings.mlr.press/v293/spiess25a.html},
}Cite STraceBERT: Source Code Retrieval using Semantic Application Traces
@inproceedings{spiess2023stracebert,
title = {{STraceBERT: Source Code Retrieval using Semantic Application Traces}},
author = {Spiess, Claudio},
year = {2023},
booktitle = {Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering},
pages = {2207--2209},
publisher = {ACM},
url = {https://doi.org/10.1145/3611643.3617852},
doi = {10.1145/3611643.3617852},
}