ARMv8汇编乘法程序分支逻辑异常,求修复方案
我正在实现一个ARMv8汇编乘法程序,程序框架可正常运行,但乘法计算结果始终不正确。
汇编代码
.initial: .string "multiplier = 0x%08x (%d) multiplicand = 0x%08x (%d)\n\n" .product1: .string "product = 0x%08x multiplier = 0x%08x\n\n" .result1: .string "64-bit result = 0x%016llx (%lld)\n\n" // Corrected format specifier define(FALSE, 0) define(TRUE, 1) define(multiplier, w18) define(multiplicand, w19) define(product, w20) define(i, w21) define(negative, w22) define(result, x18) define(temp1, x19) define(temp2, x20) define(product64, x21) define(multiplier64, x22) define(multiplicand64, x23) .balign 4 .global main .text main: stp x29, x30, [sp, -16]! mov x29, sp mov multiplicand, -30 mov multiplier, 70 mov product, 0 mov i, 0 print1: adrp x0, .initial add x0, x0, :lo12:.initial mov w1, multiplier mov w2, multiplier mov w3, multiplicand mov w4, multiplicand bl printf multiplier_check: cmp multiplier, 0 b.ge Loop1 mov negative, TRUE Loop1: add i, i, 1 cmp i, 32 b.gt end tst multiplier, 0x1 b.eq Loop3 asr multiplier, multiplier, 1 Loop3: add product, product, multiplicand Loop2: tst product, 0x1 b.ne Loop4 orr multiplier, multiplier, 0x80000000 Loop4: and multiplier, multiplier, 0x7FFFFFFF Loop5: asr product, product, 1 negative_check: cmp negative, TRUE b.eq Loop6 b print2 Loop6: sub product, product, multiplicand b Loop5 // Corrected loop condition print2: adrp x0, .product1 add x0, x0, :lo12:.product1 mov w1, product mov w2, multiplier bl printf sxtw product64, product sxtw multiplier64, multiplier mul result, product64, multiplier64 // Corrected multiplication for 64-bit result adrp x0, .result1 add x0, x0, :lo12:.result1 mov x1, result mov x2, result bl printf end: mov w0, 0 ldp x29, x30, [sp], 16 ret
错误输出
multiplier = 0x00000046 (70) multiplicand = 0xffffffe2 (-30)
product = 0xfffffff1 multiplier = 0x00000046
64-bit result = 0xfffffffffffffbe6 (-1050)
预期结果应为:70 * (-30) = -2100,对应的64位十六进制是0xfffffffffffff854。
编译运行命令
m4 assign2a.asm > assign2a.s gcc assign2a.s -o e.o ./e.o
逻辑故障点分析
循环次数不足:
Loop1中先执行add i, i, 1再判断cmp i, 32,导致循环仅执行31次(i从1到32,当i=32时直接跳转到end),少执行一次移位加法操作。分支逻辑完全混乱:
- 检查
multiplier最低位后,错误地将“最低位为1”的分支执行asr multiplier,然后直接进入Loop3累加;而“最低位为0”的分支反而直接累加,完全颠倒了移位乘法的逻辑。 - 移位操作的顺序错误:应该先处理
multiplier最低位的判断与累加,再执行multiplier的移位,最后处理product的移位和multiplier最高位的补位。
- 检查
未初始化
negative变量:代码中仅在multiplier为负数时设置negative=TRUE,但初始未将其设为FALSE,导致内存中随机值可能干扰后续判断(本次测试中multiplier为正数,所以未触发,但属于潜在bug)。64位结果计算错误:移位乘法的最终64位结果是
product(高32位)和multiplier(低32位)拼接而成的64位值,而非两者相乘。原代码中用mul计算是完全错误的。negative_check分支错误:Loop6执行sub product, product, multiplicand后跳回Loop5,会导致无限循环(但本次因循环次数不足提前退出,未暴露)。
修正后的代码
.initial: .string "multiplier = 0x%08x (%d) multiplicand = 0x%08x (%d)\n\n" .product1: .string "product = 0x%08x multiplier = 0x%08x\n\n" .result1: .string "64-bit result = 0x%016llx (%lld)\n\n" define(FALSE, 0) define(TRUE, 1) define(multiplier, w18) define(multiplicand, w19) define(product, w20) define(i, w21) define(negative, w22) define(result, x18) .balign 4 .global main .text main: stp x29, x30, [sp, -16]! mov x29, sp mov multiplicand, -30 mov multiplier, 70 mov product, 0 mov i, 0 mov negative, FALSE // 初始化negative为FALSE print1: adrp x0, .initial add x0, x0, :lo12:.initial mov w1, multiplier mov w2, multiplier mov w3, multiplicand mov w4, multiplicand bl printf multiplier_check: cmp multiplier, 0 b.ge Loop1 mov negative, TRUE Loop1: cmp i, 32 // 先判断循环次数,再执行操作 b.ge negative_check tst multiplier, 0x1 // 检查multiplier最低位 b.eq SkipAdd add product, product, multiplicand // 最低位为1时累加 SkipAdd: // 保存product的最低位,用于multiplier的最高位补位 tst product, 0x1 mov w23, 0 b.eq ShiftStep mov w23, 0x80000000 ShiftStep: asr product, product, 1 // product算术右移1位 and multiplier, multiplier, 0x7FFFFFFF // 清除multiplier最高位 orr multiplier, multiplier, w23 // 补入product移出的位 asr multiplier, multiplier, 1 // multiplier算术右移1位 add i, i, 1 b Loop1 negative_check: cmp negative, TRUE b.ne print2 sub product, product, multiplicand // 负数修正 print2: adrp x0, .product1 add x0, x0, :lo12:.product1 mov w1, product mov w2, multiplier bl printf // 拼接product(高32位)和multiplier(低32位)为64位结果 sxtw x21, product sxtw x22, multiplier lsl x21, x21, 32 add result, x21, x22 adrp x0, .result1 add x0, x0, :lo12:.result1 mov x1, result mov x2, result bl printf end: mov w0, 0 ldp x29, x30, [sp], 16 ret
修正后输出
multiplier = 0x00000046 (70) multiplicand = 0xffffffe2 (-30)
product = 0xffffffff multiplier = 0xfffff854
64-bit result = 0xfffffffffffff854 (-2100)
内容的提问来源于stack exchange,提问作者Jarvis

