x86汇编实现字符串压缩时出现Segmentation Fault问题排查求助
字符串压缩C++转x86汇编段错误排查
我正在将以下C字符串压缩代码改写为x86汇编版本,代码包含字符串比较与压缩的循环逻辑,但汇编代码运行时出现Segmentation Fault,怀疑是比较循环导致问题。以下是C原代码和我编写的x86汇编代码,恳请帮忙定位问题所在。
C++原代码
#include <string> #include <vector> using namespace std; int Min(int A, int B) { return A < B ? A : B; } int solution(string s) { int answer = s.length(); for (int Len = 1; Len <= s.length() / 2; Len++) { string Result = ""; string Temp = s.substr(0, Len); int Cnt = 1; int i; for (i = Len; i <= s.length(); i = i + Len) { if (Temp == s.substr(i, Len)) Cnt++; else { if (Cnt == 1) Result += Temp; else { Result += to_string(Cnt); Result += Temp; } Temp = s.substr(i, Len); Cnt = 1; } } if (i > s.length()) { if (Cnt == 1) Result += Temp; else { Result += to_string(Cnt); Result += Temp; } } answer = Min(answer, Result.length()); } return answer; }
编写的x86汇编代码
compress: mov rcx, len cmp rcx, 1 je len1 shr rcx, 1 ; Repeat for half the length of the string, rcx = len/2 mov rbx, 1 ; rbx = i, compression unit for the string mov rdx, len ; rdx stores the length of the array compress_loop: cmp rbx,rcx ; If i > len/2, exit loop jg done push rbx push rcx push rdx lea rsi, [array] lea rdi, [temp] lea r8, [result] call substr ; Copy i elements of array to temp xor r10, r10 inc r10 ; count = 1 mov r11, rbx ; r11 = j call compare push rsi call find_len pop rsi push rax call clean pop rax pop rdx pop rcx pop rbx inc rbx ; i++ cmp rdx, rax ; Compare rax with the length of the array ja find_answer ; Store the smaller value in rdx jmp compress_loop compare: cmp r11, rdx ; If j > len, exit loop jg compare_done mov r12, rsi add r12, r11 ; r12 = array[j] mov r13, rdi ; r13 = temp jmp compare_loop1 compare_loop1: mov al, byte[r13] mov bl, byte[r12] cmp al, 0 ; Until temp is finished, if al = bl, count + 1 je count_plus cmp al, bl je compare_loop2 jmp new compare_loop2: inc r13 inc r12 jmp compare_loop1 compare_done : ret count_plus: inc r10 ; count + 1 add r11, rbx ; j = j + i jmp compare new : cmp r10, 1 je copy_count1 jmp copy_string copy_count1: ; If count = 1, store temp in result, and store a new string in temp push rcx push rsi ; array push rdi ; temp mov rsi, rdi mov rdi, r8 mov rcx, rbx rep movsb ; Store temp in result pop rdi ; rdi = temp call clean_temp ; Initialize temp to 0 pop rsi add rsi, r11 ; rsi = array[j] rep movsb ; Store a new string in temp pop rcx sub rdi, rbx ; rdi = temp[0] lea rsi, [array] ; rsi = array[0] mov r10, 1 ; count = 1 add r11, rbx ; j = j + i jmp compare copy_string : ; If count != 1, store temp + count in result push rcx push rsi push rdi push rax mov rsi, rdi mov rdi, r8 mov rcx, rbx rep movsb ; Store temp in result mov rax, r10 call copy_count ; Add count to result pop rax pop rdi call clean_temp pop rsi add rsi, r11 ; rsi = array[j] rep movsb ; Store a new string in temp pop rcx sub rdi, rbx ; rdi = temp[0] lea rsi, [array] ; rsi = array[0] mov r10, 1 ; count = 1 add r11, rbx ; j = j + i jmp compare
问题定位分析
- 寄存器非法修改:在
compare_loop1中,使用bl寄存器读取array[r12]的字符,但rbx是外层循环的压缩单元长度(Len),修改bl会直接破坏这个关键变量,导致后续add r11, rbx等计算错误,甚至访问超出数组范围的内存。 - 字符串越界访问:原C++中
s.substr(i, Len)会自动处理i + Len超过字符串长度的情况,只取剩余字符;但汇编里直接从array[j]开始读取Len个字符,当j + Len > len时,会访问数组外的非法内存,触发段错误。 - rep movsb使用错误:在
copy_count1和copy_string中,执行rep movsb时:rcx的值可能未正确恢复(比如copy_count1中push的rcx是外层循环的len/2,并非要复制的长度);rdi(result的指针)没有维护偏移,多次执行rep movsb会覆盖之前写入的内容,或者超出result缓冲区的范围。
- 调用约定违反:x86-64系统V调用约定中,rbx、rbp、r12-r15是非易失性寄存器,需要在函数调用前保存并在调用后恢复。但代码中调用
substr、clean_temp等函数时,没有保存这些寄存器,导致它们的值被破坏,后续逻辑混乱。 - 循环收尾逻辑缺失:原C++中当
i > s.length()时,会把最后一组Temp和Cnt追加到Result中,但汇编的compare_done直接返回,没有处理这部分逻辑,导致Result不完整,后续计算长度时可能访问非法内存。
内容的提问来源于stack exchange,提问作者user25004318
相关产品推荐
相关产品推荐

