Python subprocess.run捕获输出时偶数字节返回触发malloc崩溃问题排查
在Python 3.10.12中使用subprocess.run执行外部可执行文件communicator,该程序读取传感器数值,会打印执行时间和二进制格式的传感器读数。使用的代码如下:
command = ["communicator","run","-d","9"] result = subprocess.run(command, shell=True, capture_output=True) print(f"STDOUT : {result.stdout}.") print(f"STDERR : {result.stderr}.")
其中-d参数指定返回的字节数:
- 当返回奇数字节时,运行正常,输出有效:
示例1:
STDOUT : b'Execution time: 626 ms\n 0 1 2 3 4 5 6 7 8 9 a b c d e f\n0000: 5c 5c 28 "\(\xb3"\n'.
STDERR : b'Result : 0x0\n'.示例2:
STDOUT : b'Execution time: 540 ms\n 0 1 2 3 4 5 6 7 8 9 a b c d e f\n0000: 5c 1c 28 10 41 "\.(.A\x7f"\n'.
STDERR : b'Result: 0x0\n' - 当返回偶数字节(如4、6、8字节)时,程序崩溃,stderr出现malloc错误:
STDOUT : b''.
STDERR : b'malloc(): corrupted top size\nAborted (core dumped)\n'.
但直接在Linux终端运行该命令,或设置capture_output=False时,无论字节数奇偶都能正常输出。需要排查malloc崩溃原因,并找到能正常捕获所有情况输出的方法。
一、先移除shell=True的影响
当前代码使用shell=True会启动shell进程包裹目标命令,可能引入额外的输出处理逻辑,尤其是二进制输出场景下容易出现异常。先改为直接执行程序:
command = ["communicator","run","-d","9"] result = subprocess.run(command, shell=False, capture_output=True, text=False) print(f"STDOUT : {result.stdout}.") print(f"STDERR : {result.stderr}.")
二、定位malloc崩溃的根本原因
malloc错误本质是communicator程序自身的内存问题,仅在subprocess捕获输出时触发:
- 生成core dump分析:在终端执行
ulimit -c unlimited开启core dump生成,再用subprocess运行触发崩溃的命令,生成core文件后,通过gdb communicator core查看崩溃栈,定位具体的内存错误位置。 - 检查程序输出逻辑:程序输出二进制内容时,可能存在字符串终止符处理异常(比如偶数字节时,最后一个字节触发了shell或subprocess缓冲区的错误处理逻辑)。
三、绕开subprocess内存缓冲的临时方案
如果暂时无法修复communicator的内存问题,可以通过文件重定向的方式捕获输出,避免内存缓冲触发崩溃:
import tempfile import os with tempfile.NamedTemporaryFile(mode='wb', delete=False) as out_f, \ tempfile.NamedTemporaryFile(mode='wb', delete=False) as err_f: command = ["communicator","run","-d","8"] subprocess.run(command, shell=False, stdout=out_f.fileno(), stderr=err_f.fileno()) out_f.close() err_f.close() with open(out_f.name, 'rb') as f: stdout_content = f.read() with open(err_f.name, 'rb') as f: stderr_content = f.read() print(f"STDOUT : {stdout_content}.") print(f"STDERR : {stderr_content}.") os.unlink(out_f.name) os.unlink(err_f.name)
四、调整subprocess缓冲区参数
尝试修改bufsize参数,避免默认全缓冲的影响:
# 无缓冲模式 result = subprocess.run(command, shell=False, capture_output=True, bufsize=0) # 或逐块读取输出 from subprocess import Popen, PIPE with Popen(command, shell=False, stdout=PIPE, stderr=PIPE, bufsize=1) as proc: stdout, stderr = proc.communicate() print(f"STDOUT : {stdout}.") print(f"STDERR : {stderr}.")
内容的提问来源于stack exchange,提问作者user3777923

