Beyond the Leaderboard: Multi-Dimensional Evaluation of Dense and Mixture-of-Experts Models for Automated Program Repair
Anvi Kalpesh Shah, Umamaheswara Sharma B
https://t.co/9Rlkfa0eOE [ππ.ππ΄ ππ.π»πΆ]
Acquiring and Verifying Repository Norms for Coding Agents
Kaifeng He, Xiaojun Zhang, Zhenxi Chen, Christoph Treude, Mingwei Liu
https://t.co/7xq3WFlNP8 [ππ.ππ΄]
A Validated Dataset and Benchmark for Coherent Multi-Diagram SysML Models
Ardalan Aryashad, Yan Jin
https://t.co/ulUaNDgTTW [ππ.ππ΄ ππ.π°πΈ]
Catching Developers in the Flow: Low-Latency Agentic Program Repair at Google Scale
Celal Ziftci, Spencer Greene, Ray Liu, β¦
https://t.co/jTz0Iiyt9C [ππ.ππ΄ ππ.π°πΈ]
π¬Accepted at the 41st IEEE/ACM International Conference on Automated Software Engineering (ASE 2026)
SAFESHIELD: A Decision-Organization Framework for Deployment-Time Safety of Small Language Models
Xingru Zhou, Luis Sentis, Aarti Choudhary
https://t.co/RJMjd3x1aO [ππ.ππ΄ ππ.π°πΈ ππ.π²π]
π¬Accepted to the Application Track of IEEE TPS 2026
Choosing an energy-efficient software architecture for building system diagnostic support
Roxane Koitz-Hristov, Franz Wotawa
https://t.co/Tsz9v6HE0k [ππ.ππ΄ ππ.π°πΈ]
LLM Non-Determinism in Deterministic Processes: A Controlled Variation Space for Constraining the Propagation of Non-Determinism
Paul Darius Mandl, Peter Mandl, Martin HΓ€usl
https://t.co/vMOc64HDJ9 [ππ.ππ΄]
Where Does the Budget Go? Structural Concentration in Learning-Based Mutant Selection
Taha Rostami, Jeongju Sohn, Mike Papadakis
https://t.co/3g3EpPJzjt [ππ.ππ΄]
Let the Agent Do It? How Software Practitioners Understand and Make Permission Decisions in Agentic AI Assistants
Larissa Salerno, Haoyu Gao, Gregory Gay, Alexander Serebrenik, Philipp Leitner
https://t.co/SfSpMqvoWA [ππ.ππ΄]
AgentSpy: Making AI Agent Behavior Observable
Christoph BΓΌhler, Matteo Biagiola, Luca Di Grazia, Guido Salvaneschi
https://t.co/EbtONNkWBk [ππ.ππ΄ ππ.π°πΈ]
From Trust-by-Design to Continuous Trust in the Era of Artificial Intelligence
Peter Fettke, Wolfgang Reisig
https://t.co/JCodF7t5Ao [ππ.ππ΄]
Mechanizing the User's Eye: Pre-Registered Deployment of a Sabotage-Validated Fail-Plausible Observer in a Production LLM Agent Runtime
Wei Wu
https://t.co/ClqWSOmGxy [ππ.ππ΄]
π¬Code: https://t.co/Sp33IywKhm