Survival of LLM repairs across the mobile runtime boundary
Determine whether repairs produced by repository-level LLM agents survive the complete mobile build–install–launch–test boundary, rather than being misclassified because missing SDKs, offline devices, or pre-assertion application crashes prevent the intended behavior test from executing.
References
It remains unclear whether their repairs survive the mobile build--install--launch--test boundary, where a missing SDK, offline device, or pre-assertion crash can be mistaken for a program failure.
— AppEval: A Unified Benchmark for LLM-Based Mobile Application Repair in ArkTS, Swift, and Kotlin
(2608.18588 - Xie et al., 19 Aug 2026) in Abstract