You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

ELF/Linux环境下如何中断栈/调用帧信息链?

关于x86_64 Linux下修改栈布局导致DWARF CFI unwind崩溃的问题

我正尝试在ELF/Linux环境下实现一项特殊需求:中断DWARF异常处理信息中的CFI(调用帧信息),以及帧间rbp与rsp的关联。核心目的是在线程控制流的特定节点后,实现一种结合单向尾调用和yield的调用延续机制——清理栈后回到栈顶,以便在延续点重新执行。

以下是原理性实现思路,只要不启用修改栈的代码就能正常运行:

/* x86_64 SysV:
 * rdi, rsi, rdx, rcx, r8, r9, xmm0-xmm7
 */
__asm {
    mov rax, TCB
    mov rax, qword ptr [rax] OSThreadControlBlock.StartFn;
    call rax;
    mov rax, 0; // end of stack
    //push rax;
    //push rax;
    //push rbx; // last "real" frame
    //push rbp;
    //mov rbp, rsp;
    //push rbx; // make the call
    mov rdi, RL;
    lea rax, qword ptr __OS_RUNLOOP_START__;
    call rax;
    // trap if it returns
    //int 3;
}

我了解SP/BP寄存器的基本原理,且已指定使用-fno-omit-frame-pointer编译选项。但折腾了数小时仍未成功,我到底遗漏了什么?似乎任何栈布局修改——哪怕是调用前简单的push操作(保持栈对齐)——都会引发连锁崩溃,自定义信号捕获到的错误如下:

Received fatal signal: Segmentation fault (11) [thread: 10298 ctl-thrd]
* Unknown error at address 0x0
Regs:
%rip=0x00000000003E2D91 %rbp=0x00007F820A547EA8 %rsp=0x00007F820A547DE8
%rax=0x00007F820A547DE8 %rbx=0x00007F820A547F38 %rdi=0x00000000002121E1
%rsi=0x000000000000007B %rcx=0x000000000000000A %r8=0x0000000000000900
%r9=0x00007F820A5490C0

当前使用的ABI是x86_64 Linux平台下的libc++/libc++abi,基于LLVM/Clang 6.0.X工具链。我尝试了各种方法,上述内联汇编是MS扩展语法,已多次通过反汇编确认生成的代码正常。我推测这是CFI与基于帧指针的机制之间的冲突,但对x86_64细节不够熟悉,无法确定问题所在。我知道unwind过程应由哨兵(最后一帧的SP/FP为空)终止,但目前连调试器都被干扰,完全陷入困境。

只要修改栈,哪怕之后恢复原状,程序就会彻底崩溃。由于最后一次调用无需常规返回,汇编块外的寄存器破坏无关紧要。我还注意到这似乎与TLV有关,但不确定NPTL的配置如何影响它。

恳请各位提供建议或帮助。


编辑补充:

Valgrind的以下注释似乎能解释问题原因:

/* NB 9 Sept 07. There is a nasty kludge here in all these CALL_FN_ macros. In order not to trash the stack redzone, we need to drop %rsp by 128 before the hidden call, and restore afterwards. The nastyness is that it is only by luck that the stack still appears to be unwindable during the hidden call - since then the behaviour of any routine using this macro does not match what the CFI data says. Sigh. Why is this important? Imagine that a wrapper has a stack allocated local, and passes to the hidden call, a pointer to it. Because gcc does not know about the hidden call, it may allocate that local in the redzone. Unfortunately the hidden call may then trash it before it comes to use it. So we must step clear of the redzone, for the duration of the hidden call, to make it safe. Probably the same problem afflicts the other redzone-style ABIs too (ppc64-linux, ppc32-aix5, ppc64-aix5); but for those, the stack is self describing (none of this CFI nonsense) so at least messing with the stack pointer doesn't give a danger of non-unwindable stack. */

内容的提问来源于stack exchange,提问作者krb

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 07:16:30