| A Benchmark for Language Models in Real-World System Building |
Weilin聽Jin, Chenyu聽Zhao, Zeshun聽Huang, Chaoyun聽Zhang, Qingwei聽Lin, Chetan聽Bansal, Saravan聽Rajmohan, Shenglin聽Zhang, Yongqian聽Sun, Dan聽Pei, Yifan聽Wu, Tong聽Jia, Ying聽Li, Zhonghai聽Wu, Minghua聽Ma |
| A Spec-Driven Workflow for AI-Assisted Domain-Driven Development: Insights from Practice 馃摑 |
Jefferson de Barros聽Santos |
| Achieving Productivity Gains with AI-based IDE features: A Journey at Google |
Maxim聽Tabachnyk, Xu聽Shu, Alexander聽Fr枚mmgen, Pavel聽Sychev, Vahid聽Meimand, Ilia聽Krets, Stanislav聽Pyatykh, Abner聽Araujo, Kristof聽Molnar, Satish聽Chandra |
| An Automated Methodology for Generating Labeled Datasets of Semantic Errors in Code |
Mahmoud聽Kassem, Francisco聽Ribeiro, Sarah聽Nadi |
| An Empirical Study of C to Rust Translation using Local Large-Language Models |
Nathan聽Rutherford, Dan聽O'Keeffe |
| An Initial Exploration of Contrastive Prompt Tuning to Generate Energy-Efficient Code |
Sophie聽Weidmann, Fernando聽Castor |
| Benchmarking LLM Commit Message Generation through a Developer-centric Pairwise Preference Framework |
Lucas聽Aguiar, Matheus聽Freitas, Matheus聽Paixao, Rafael聽Carmo |
| Code Roulette: How Prompt Variability Affects LLM Code Generation |
Andrei聽Paleyes, Diana聽Robinson, Radzim聽Sendyka, Christian聽Cabrera, Neil D.聽Lawrence |
| Code vs Serialized AST Inputs for LLM-Based Code Summarization: An Empirical Study |
Shijia聽Dong, Haoruo聽Zhao, Paul聽Harvey |
| ContextPilot: Code Context Engineering with Memory-Augmented Exploration Agents 馃摑 |
Shuzheng聽Gao, Chaozheng聽Wang, Shuqing聽Li, Yun聽Peng, Michael R.聽Lyu |
| Continuous Benchmark Generation for Evaluating Enterprise-scale LLM Agents 馃摑 |
Divyanshu聽Saxena, Rishikesh聽Maurya, Xiaoxuan聽Ou, Gagan聽Somashekar, Shachee Mishra聽Gupta, Arun聽Iyer, Yu聽Kang, Chetan聽Bansal, Aditya聽Akella, Saravan聽Rajmohan |
| CP-Agent: Agentic Constraint Programming |
Stefan聽Szeider |
| Diverse LLMs vs. Vulnerabilities: Who Detects and Fixes Them Better? |
Arastoo聽Zibaeirad, Marco聽Vieira |
| Do LLMs Dream of Energy-Efficient Code? |
Antimo Di聽Bernardo, Gianluca聽Capozzi, Pasquale De聽Rosa, Daniele Cono聽D'Elia, Leonardo聽Querzoni, Giuseppe Antonio Di聽Luna, Valerio聽Schiavoni |
| English or Chinese? Investigating the Impact of Prompt Language on Large Language Models for Code Summarization 馃摑 |
Yijia聽Tang, Zhiqiu聽Huang, Jian聽Xie, Yaoshen聽Yu, Bowei聽Xia, Enya聽Shen, Yukun聽Cao |
| Evaluating LLMs-Driven Java Code Refactoring from a Developer鈥檚 Perspective 馃挰 |
Javel聽Freitas, Guilherme聽Pereira, Lara聽Lima, Caio Rian de聽Sousa, Edivar聽Filho, Jos茅 Cezar de Souza聽Filho, Paulo Henrique聽Maia, Carla聽Bezerra |
| Learning Functional Equivalence via Supervised Contrastive Code-Problem Alignment |
Siu Wun聽Cheung, Harshitha聽Menon |
| LLM-Driven SQL Remediation: Towards Safe and Explainable Code for Automated Schema Refactoring |
Antony聽Medeiros, Claudio聽Cavalcante, Nicolaas聽Ruberg, Sergio聽Lifschitz |
| LLM-Powered On-Demand Test Suites in Self-Graded Student Programming Assignments 馃挰 |
Chang聽Liu |
| MAsFL: Data-Secure, Efficient and Accurate Fault Localization with Multi-Agent Small Language Models |
DUONG PHAM聽DUC, HIROSHI聽SATO, MASAO聽KUBO |
| Multi-task Code LLMs: Data Mix or Model Merge? |
Mingzhi聽Zhu, Michele聽Merler, Stacy聽Patterson, Raju聽Pavuluri, Rahul聽Krishna, Boris聽Sobolev |
| Natural Language Summarization Enables Multi-Repository Bug Localization by LLMs in Microservice Architectures |
Amirkia Rafiei聽Oskooei, S. Selcan聽Yukcu, Mehmet Cevheri聽Bozoglan, Mehmet S.聽Aktas |
| RAG Against the Machine: Zero-Shot Software Vulnerabilities Classification using LLMs |
Edvin聽Nordqvist, Changjie聽Wang, Simone聽Ferlin, Mariano聽Scazzariello, Marco聽Chiesa |
| RubberDuckBench: A Benchmark for AI Coding Assistants |
Elizabeth聽Dinella, Ferida聽Mohammed, Fatma聽Ayad, Satish聽Chandra, Petros聽Maniatis |
| SecRepoBench: Benchmarking Code Agents for Secure Code Completion in Real-World Repositories |
Chihao聽Shen, Connor聽Dilgren, Purva聽Chiniya, Luke聽Griffith, Yu聽Ding, Yizheng聽Chen |
| Statistical Independence Aware Caching for LLM Workflows |
Yihan聽Dai, Dimitrios Stamatios聽Bouras, Haoxiang聽Jia, Sergey聽Mechtaev |
| The Hidden DNA of LLM-Generated JavaScript: Structural Patterns Enable High-Accuracy Authorship Attribution |
Norbert聽Tihanyi, Bilel聽Cherif, Mohamed Amine聽Ferrag, Richard A.聽Dubniczky, Tamas聽Bisztray |
| Towards Improving in-IDE Code Completion for Driver Development |
Batuhan Raif聽Karagoz, Mahesh聽Jayasankar, Saurabh聽Bodhe, Subhayan聽Roy, Lejin聽Varghese, Max聽Kiehn, Yonas聽Bedasso |
| Towards LLM-guided Semantic Validation of Autonomous Driving Safety Policies 馃摑 |
Qingzhao聽Zhang, Z. Morley聽Mao |
| TritonForge: Profiling-Guided Framework for Automated Triton Kernel Optimization |
Haonan聽Li, Keyu聽Man, Partha聽Kanuparthy, Hanning聽Chen, Wei聽Sun, Sreen聽Tallam, Chenguang聽Zhu, Kevin聽Zhu, Zhiyun聽Qian |
| Usage, Effects and Requirements for AI Coding Assistants in the Enterprise: An Empirical Study |
Michele聽Merler, Rangeet聽Pan, Rahul聽Krishna, Tin Kam聽Ho, Raju聽Pavuluri, Maja聽Vukovic |