如何从Android TV应用的Runtime aborting崩溃中定位问题
Android TV应用高负载场景下周期性崩溃定位思路
应用崩溃特征:
- 多数时间运行正常
- 仅部分设备出现该崩溃
- 仅在高负载场景下触发
崩溃线程转储片段
Cmdline: com.mydomain.myapp pid: 28890, tid: 3124, name: CodecLooper >>> com.mydomain.myapp <<< Davey! duration=754ms; Flags=0, FrameTimelineVsyncId=15812495, IntendedVsync=146790515172662, Vsync=146790880252007, InputEventId=0, HandleInputStart=146790883087631, AnimationStart=146790883088506, PerformTraversalsStart=146790883617881, DrawStart=146790899276008, FrameDeadline=146790546918692, FrameInterval=146790883080923, FrameStartTime=15873015, SyncQueued=146790906193467, SyncStart=146791117319030, IssueDrawCommandsStart=146791122452364, SwapBuffers=146791474395774, FrameCompleted=146791480789108, DequeueBufferDuration=340854242, QueueBufferDuration=5631208, GpuCompleted=146791476995983, SwapBuffersCompleted=146791480789108, DisplayPresentTime=0, runtime.cc:669] Runtime aborting... runtime.cc:669] Dumping all threads without mutator lock held runtime.cc:669] All threads: runtime.cc:669] DALVIK THREADS (260): runtime.cc:669] "pool-139-thread-4" prio=5 tid=402 Runnable runtime.cc:669] | group="" sCount=0 ucsCount=0 flags=0 obj=0x162037f0 self=0xb4000075be884a40 runtime.cc:669] | sysTid=3110 nice=0 cgrp=default sched=0/0 handle=0x721b9cacb0 runtime.cc:669] | state=R schedstat=( 276302280598 789549905823 2816620 ) utm=22266 stm=5363 core=4 HZ=100 runtime.cc:669] | stack=0x721b8c7000-0x721b8c9000 stackSize=1039KB runtime.cc:669] | held mutexes= "abort lock" "mutator lock"(shared held) runtime.cc:669] native: #00 pc 000000000055f850 /apex/com.android.art/lib64/libart.so (art::DumpNativeStack(std::__1::basic_ostream<char, std::__1::char_traits<char> >&, int, BacktraceMap*, char const*, art::ArtMethod*, void*, bool)+140) runtime.cc:669] native: #01 pc 0000000000676270 /apex/com.android.art/lib64/libart.so (art::Thread::DumpStack(std::__1::basic_ostream<char, std::__1::char_traits<char> >&, bool, BacktraceMap*, bool) const+360) runtime.cc:669] native: #02 pc 0000000000693f4c /apex/com.android.art/lib64/libart.so (art::DumpCheckpoint::Run(art::Thread*)+920) runtime.cc:669] native: #03 pc 000000000068da70 /apex/com.android.art/lib64/libart.so (art::ThreadList::RunCheckpoint(art::Closure*, art::Closure*)+520) runtime.cc:669] native: #04 pc 000000000068cc84 /apex/com.android.art/lib64/libart.so (art::ThreadList::Dump(std::__1::basic_ostream<char, std::__1::char_traits<char> >&, bool)+1464) runtime.cc:669] native: #05 pc 0000000000626d68 /apex/com.android.art/lib64/libart.so (art::Runtime::Abort(char const*)+2164) runtime.cc:669] native: #06 pc 000000000001595c /system/lib64/libbase.so (android::base::SetAborter(std::__1::function<void (char const*)>&&)::$_3::__invoke(char const*)+76) runtime.cc:669] native: #07 pc 0000000000006dc8 /system/lib64/liblog.so (__android_log_assert+308) runtime.cc:669] native: #08 pc 000000000001ad34 /system/lib64/libstagefright_foundation.so (android::ALooperRoster::registerHandler(android::sp<android::ALooper> const&, android::sp<android::AHandler> const&)+796) runtime.cc:669] native: #09 pc 000000000001966c /system/lib64/libstagefright_foundation.so (android::ALooper::registerHandler(android::sp<android::AHandler> const&)+136) runtime.cc:669] native: #10 pc 000000000010c470 /system/lib64/libstagefright.so (android::MediaCodec::init(android::AString const&)+1556) runtime.cc:669] native: #11 pc 0000000000049130 /system/lib64/libmedia_jni.so (android_media_MediaCodec_reset(_JNIEnv*, _jobject*)+328) runtime.cc:669] at android.media.MediaCodec.native_reset(Native method) runtime.cc:669] at android.media.MediaCodec.reset(MediaCodec.java:1987)
崩溃前1秒的BLAST Consumer日志
#17(BLAST Consumer)17](id:70da00000011,api:3,p:28890,c:28890) detachBuffer: slot 39 is not owned by the producer (state = FREE) #17(BLAST Consumer)17](id:70da00000011,api:3,p:28890,c:28890) detachBuffer: slot 40 is not owned by the producer (state = FREE) #17(BLAST Consumer)17](id:70da00000011,api:3,p:28890,c:28890) detachBuffer: slot 41 is not owned by the producer (state = FREE) #17(BLAST Consumer)17](id:70da00000011,api:3,p:28890,c:28890) detachBuffer: slot 42 is not owned by the producer (state = FREE) #17(BLAST Consumer)17](id:70da00000011,api:3,p:28890,c:28890) detachBuffer: slot 43 is not owned by the producer (state = FREE) ... #17(BLAST Consumer)17](id:70da00000011,api:3,p:28890,c:28890) detachBuffer: slot 63 is not owned by the producer (state = FREE)
定位思路
1. 聚焦MediaCodec重置流程的异常
从崩溃栈可见,崩溃触发于MediaCodec.reset()调用,底层ALooperRoster.registerHandler触发断言失败。核心排查点:
- 确认MediaCodec对象状态:是否存在已释放对象被重复调用reset的情况
- 检查reset调用时机:是否在高负载下线程调度延迟,导致Handler注册时Looper或Handler状态不一致
2. 关联BLAST缓冲异常与崩溃的因果
崩溃前大量出现的缓冲区状态异常日志,说明应用的MediaCodec输出Surface与BLAST消费者端的缓冲状态不一致。高负载下:
- 应用侧MediaCodec因卡顿未及时处理缓冲区,导致BLAST尝试释放已处于FREE状态的缓冲,触发系统层警告
- 这种缓冲异常可能触发MediaCodec的错误处理逻辑(如自动触发reset),而reset过程中恰好遇到底层断言失败,引发致命崩溃
3. 针对设备特异性的排查
仅部分设备出现问题,需对比差异:
- 设备的Android版本、芯片厂商定制的MediaCodec驱动:不同厂商的Codec实现可能存在兼容性bug
- 设备的SurfaceFlinger/BLAST版本:旧版本或定制ROM可能存在缓冲管理逻辑缺陷
- 设备的硬件资源阈值:高负载下内存/CPU不足,导致线程调度异常,放大了潜在的状态不一致问题
4. 高负载场景下的复现与验证
- 模拟高负载:同时播放多视频、启动后台进程占用资源,复现崩溃场景
- 补充日志:在MediaCodec的创建、reset、释放流程中添加详细状态日志,记录调用线程、对象生命周期阶段
- 采集系统日志:使用
logcat -b system捕获MediaCodec、SurfaceFlinger的系统层日志,追踪崩溃前的Codec状态变化
5. 代码层面的优化验证
- 规范MediaCodec生命周期:确保reset、release操作在指定线程执行,避免多线程并发调用导致的状态混乱
- 增加异常防护:在
MediaCodec.reset()调用处添加try-catch,捕获异常并记录状态,避免触发致命崩溃 - 优化缓冲交互:确保MediaCodec与Surface的缓冲区操作符合系统规范,减少状态不一致的概率
内容的提问来源于stack exchange,提问作者Hong
相关产品推荐
相关产品推荐

