Publications

, ,
How Robustly Do LLMs Understand Execution Semantics?
AIware 2026, 2026
We study the robustness of LLM program-output prediction under code transformations and input perturbations, including inputs that raise exceptions. The results expose failures hidden by performance on the original benchmark.
, , , , , ,
Does In-IDE Calibration of Large Language Models work at Scale?
FSE 2026, 2026
Collaboration with JetBrains Research Prior work suggests calibration can improve alignment, but at-scale evidence is limited. In this work, we investigate the feasibility of applying calibration of code models to an in-IDE context. We study two aspects of the problem: (1) the technical method for implementing confidence calibration and improving the reliability of code generation models, and (2) the human-centered design principles for effectively communicating reliability signal to developers.
, , , ,
Model See, Model Do? Exposure-Aware Evaluation of Bug-vs-Fix Preference in Code LLMs
MSR 2026, 2026
We study how training-data exposure to buggy and fixed code affects code LLM preferences and generations, using membership testing to distinguish memorization from generalization.
, , ,
On LLMs’ Internal Representation of Code Correctness
ICSE 2026, 2026
Inspired by findings that LLMs internally encode concepts like truthfulness, this paper explores if LLMs similarly represent code correctness. Specifically, we identify a correctness representation inside LLMs by contrasting the hidden states between pairs of correct and incorrect code for the same programming tasks. By experimenting on four LLMs, we show that exploiting this extracted correctness representation outperforms standard log-likelihood ranking, as well as verbalized model confidence. Furthermore, we explore how this internal correctness signal can be used to select higher-quality code samples, without requiring test execution.
, , , , , , , ,
Calibration and Correctness of Language Models for Code
ICSE 2025, 2025
Machine learning models are widely used but can also often be wrong. Users would benefit from a reliable indication of whether a given output from a given model should be trusted, so a rational decision can be made whether to use the output or not. In this work, we studied whether current LLMs for code are well calibrated.
, , ,
AutoPDL: Automatic Prompt Optimization for LLM Agents
AutoML 2025, 2025
Work done at IBM Research This paper proposes AutoPDL, an automated approach to discover good LLM agent configurations. Our method frames this as a structured AutoML problem over a combinatorial space of agentic and non-agentic prompting patterns and demonstrations, using successive halving to efficiently navigate this space. We introduce a library implementing common prompting patterns using the PDL prompt programming language. AutoPDL solutions are human-readable, editable, and executable PDL programs that use this library.

STraceBERT: Source Code Retrieval using Semantic Application Traces
ESEC/FSE 2023, 2023
Novel approach that utilizes a Java dynamic analysis tool to record calls to core Java libraries, and a BERT-style model on the recorded application traces for effective method source code retrieval from a candidate set. Experiments demonstrate the effectiveness in retrieving the source code compared to existing approaches. Proposed approach offers a promising solution to the problem of code retrieval in software reverse engineering and opens up new avenues for further research in this area.