Abstract. Evaluating Large Language Models (LLMs) for code generation often reveals a significant gap between theoretical leaderboard scores and day-to-day development environments. This talk shares the findings, methodologies, and raw experiences from practical experience in running and evaluating local LLMs in the research domain.
We discuss the tools, infrastructure design, challenges faced, the successes achieved, and the lessons learned from benchmarking various LLMs for code generation tasks.
Note. Slides presented at the "seLIA: I Jornada sobre Software Libre e Inteligencia Artificial Abierta".