Large Language Models (LLMs) and agentic AI systems are rapidly transforming software engineering research, yet the community lacks consensus on how such approaches should be evaluated. This thesis systematically reviews evaluation practices in recent publications at top-tier venues.
Large Language Models (LLMs) and agentic AI systems are rapidly transforming software engineering research, yet the community lacks consensus on how such approaches should be evaluated. This thesis systematically reviews evaluation practices in recent publications at top-tier venues (e.g. ICSE, ASE, FSE) to characterise the state of practice, identify methodological weaknesses, and derive recommendations for rigorous empirical evaluation.
Publications applying LLMs and autonomous agents to software engineering tasks — code generation, program repair, testing, refactoring, requirements engineering, and beyond — have grown at an extraordinary pace. However, the empirical evaluation of these approaches raises methodological challenges that classical SE research did not face to the same degree: benchmark contamination (test data leaking into training corpora), non-determinism of model outputs, rapidly deprecating model versions that hinder reproducibility, reliance on “LLM-as-a-judge” evaluation, unclear baselines, and inconsistent reporting of costs, prompts, and hyperparameters. Whether and how the community addresses these challenges in practice remains insufficiently understood.
Goals and Tasks:
- Conduct a systematic literature review (or systematic mapping study) of LLM/agent-based SE papers published at leading venues (e.g. ICSE, ASE, FSE, ISSTA; optionally journals such as TSE, TOSEM, EMSE) within a defined time window.
- Develop a classification scheme for evaluation practices, covering, e.g., benchmark selection, baseline comparison, metrics, statistical analysis, handling of non-determinism (repetitions, variance reporting), human-subject involvement, reproducibility artefacts, and threats-to-validity reporting.
- Analyse the extracted data to identify common patterns, gaps, and methodological risks.
- Derive actionable guidelines or a checklist for evaluating LLM-based and agentic SE approaches.
- Optionally: tool-supported extraction (e.g. LLM-assisted paper screening/coding with human validation) as a meta-level contribution.
Requirements:
Interest in empirical research methods and scientific rigour; careful, systematic working style; good command of English (reading and academic writing). Programming skills are beneficial for data analysis and tool-assisted extraction but not the primary focus.
Betreuer: Stefan Klikovits und Manuel Wimmer
