解析内核模块空指针解引用生成的kdump时crash工具崩溃问题排查
排查crash工具解析aarch64 kdump时崩溃的临时方案
环境与配置
- 运行环境:qemu-system-aarch64,Linux 6.5内核,buildroot根文件系统
- kdump配置:修改
/etc/init.d/rcS添加自动生成kdump的脚本:
#hackish way of achieving things # Check if /proc/vmcore exists if [ -e "/proc/vmcore" ]; then echo "collecting core" # Get the current date and time TSTAMP=$(date +"%Y%m%d%H%M%S") # Define the filename for the kdump file FILENAME="kernel.${TSTAMP}.core.kdump" # Run makedumpfile to create the vmcore dump makedumpfile --message-level 4 -d 17,31 /proc/vmcore "${FILENAME}" reboot else echo "loading crashkernel into memory" # /proc/vmcore does not exist, so we run kexec kexec -p /Image --append="console=ttyAMA0,115200n8 root=/dev/nfs rw nfsroot=10.105.226.234:/home/naveen/nfsroot/rootfs-buildroot-arm64/,nolock,vers=4,tcp ip=10.105.226.235" fi
触发崩溃的内核模块
通过空指针解引用触发内核崩溃的模块代码:
static int __init null_deref_module_init(void) { // Pointer to an integer, initialized to NULL int *null_pointer = NULL; printk(KERN_INFO "Null dereference module loaded\n"); // Dereferencing the NULL pointer to trigger a crash printk(KERN_INFO "Triggering null pointer dereference...\n"); *null_pointer = 1; // This line will cause a null pointer dereference return 0; // This will never be reached }
问题现象
- 加载模块后,guest内核正常崩溃,kdump文件生成成功
- 使用crash 8.0.4解析kdump时工具自身崩溃,报错如下:
$ sudo crash ~/.repos/src/arm64/linux/vmlinux kernel.20240330170747.core.kdump crash 8.0.4 Copyright (C) 2002-2022 Red Hat, Inc. ...(省略版权信息)... please wait... (determining panic task)Segmentation fault
- 若将模块编译为内置模块,可得到正常崩溃信息(无kdump):
[ 63.406244] pc : null_deref_module_init+0x30/0x1000 [npdereference] [ 63.407287] lr : null_deref_module_init+0x24/0x1000 [npdereference]
调试定位的根因
调试crash工具发现,symbols.c中value_search_module_6_4函数的sp指针为空,导致空指针访问触发段错误:
Thread 1 "crash" received signal SIGSEGV, Segmentation fault. value_search_module_6_4 (value=18446603338276298752, offset=0x7ffffffface0) at symbols.c:5564 5564 if (value < sp->value) (gdb) bt #0 value_search_module_6_4 (value=18446603338276298752, offset=0x7ffffffface0) at symbols.c:5564 #1 0x0000555555812bd0 in value_to_symstr (value=18446603338276298752, buf=buf@entry=0x7fffffffb9c0 "", radix=10, radix@entry=0) at symbols.c:5872 #2 0x00005555557694a2 in display_memory (addr=<optimized out>, count=2048, flag=208, memtype=memtype@entry=1, opt=opt@entry=0x0) at memory.c:1740 #3 0x0000555555769e1f in raw_stack_dump (stackbase=<optimized out>, size=<optimized out>) at memory.c:2194 #4 0x00005555557923ff in get_active_set_panic_task () at task.c:8639 #5 0x00005555557930d2 in get_dumpfile_panic_task () at task.c:7628 #6 0x00005555557a89d3 in panic_search () at task.c:7380 #7 get_panic_context () at task.c:6267 #8 task_init () at task.c:687 #9 0x00005555557305b3 in main_loop () at main.c:787 #10 0x0000555555a64331 in captured_main (data=<optimized out>) at main.c:1284 #11 gdb_main (args=<optimized out>) at main.c:1313 #12 0x0000555555a643b0 in gdb_main_entry (argc=<optimized out>, argv=argv@entry=0x7fffffffe508) at main.c:1338 #13 0x00005555557d1ece in gdb_main_loop (argc=<optimized out>, argc@entry=3, argv=argv@entry=0x7fffffffe508) at gdb_interface.c:81 #14 0x0000555555728dfc in main (argc=3, argv=0x7fffffffe508) at main.c:720
注:模块已包含完整调试符号:
naveen@workstation:~/.repos/src/arm64/linux$ file drivers/naveen/npdereference.ko drivers/naveen/npdereference.ko: ELF 64-bit LSB relocatable, ARM aarch64, version 1 (SYSV), BuildID[sha1]=118e35b0267440ef364c551c5890ff934392fb6c, with debug_info, not stripped
临时排查与修复提示
强制加载模块符号
- 启动crash时,通过
mod -s /path/to/npdereference.ko手动指定模块符号文件路径,确认工具能否识别模块符号 - 检查makedumpfile参数
-d 17,31是否包含模块符号对应的内存页,可尝试使用-d 31(全量内存)生成kdump测试
- 启动crash时,通过
跳过panic任务自动检测
- 启动crash时添加
--no-panic-task参数,跳过自动分析panic任务的流程,直接进入交互模式,手动通过bt、ps等命令分析崩溃上下文
- 启动crash时添加
临时修补crash源码
- 在
symbols.c的value_search_module_6_4函数中,给空指针判断添加防护:if (sp && value < sp->value) - 重新编译crash工具后测试
- 在
验证版本兼容性
- 尝试使用crash 7.x版本解析kdump,确认是否为Linux 6.5与crash 8.0.4的兼容性问题
- 检查内核编译时是否开启
CONFIG_DEBUG_INFO、CONFIG_KALLSYMS、CONFIG_KALLSYMS_ALL等必要调试选项
内容的提问来源于stack exchange,提问作者InsaneCoder
相关产品推荐
相关产品推荐

