Python调用subprocess.run遇非UTF-8字符触发UnicodeDecodeError求助
问题分析
你的报错核心原因有两个:
- 设置了
universal_newlines=True(等价于text=True),会强制subprocess将命令的stdout/stderr按系统默认编码(通常是UTF-8)解码为字符串,但文件名中的0xbb字节(对应Windows-1252编码的»)不符合UTF-8规则,触发了解码失败。 - 用
bytes拼接命令字符串+shell=True的方式不仅不安全(存在shell注入风险),还容易因编码不匹配导致解析错误。
解决方案
优先采用列表传参+明确编码处理的方案,这是subprocess的最佳实践,能彻底规避编码和shell解析问题。
方案1:使用列表传参(推荐)
直接通过列表传递命令参数,subprocess会自动处理文件名的特殊字符,无需手动转义或拼接字符串:
import subprocess from subprocess import PIPE def test(): # 先确定FFP文件的编码:从文件内容的`�`来看,是Windows-1252(cp1252)编码 ffp_encoding = "cp1252" directory = "<redacted>" # 用字符串形式,确保目录名编码匹配系统环境 ffp_path = f"{directory}/<redacted>.ffp" # 用正确编码读取FFP文件,直接解析为字符串 with open(ffp_path, 'r', encoding=ffp_encoding) as ffp_open: ffp_lines = ffp_open.readlines() for line in ffp_lines: line = line.strip() if not line.startswith(';') and ':' in line: filename_part, expected_md5 = line.split(':', 1) # 构造命令列表,自动处理文件名特殊字符 cmd = [ '/usr/bin/metaflac', '--show-md5sum', f"{directory}/{filename_part}" ] # 关闭text模式,直接获取bytes输出,或指定对应编码解码 process = subprocess.run( cmd, stdout=PIPE, stderr=PIPE, encoding=ffp_encoding # 直接指定编码,避免手动解码 ) actual_md5 = process.stdout.strip() print(f"预期指纹: {expected_md5.strip()}, 实际指纹: {actual_md5}")
方案2:保留shell=True但关闭text模式
如果必须使用shell命令拼接,需关闭universal_newlines=True,直接处理bytes输出:
import subprocess from subprocess import PIPE def test(): directory = b'<redacted>' ffp_open = open(directory + b'<redacted>.ffp','rb') ffp_lines = ffp_open.readlines() for line in ffp_lines: if not line.startswith(b';') and b':' in line: txt = line.split(b':', 1) # 拼接bytes命令字符串(注意:若文件名含单引号会失效,需额外处理) ffp_cmd = b'/usr/bin/metaflac --show-md5sum \'' + directory + b'/' + txt[0].strip() + b'\'' # 去掉universal_newlines=True,默认处理bytes process = subprocess.run(ffp_cmd, stdout=PIPE, stderr=PIPE, shell=True) actual_md5 = process.stdout.strip() expected_md5 = txt[1].strip() print(f"预期指纹: {expected_md5}, 实际指纹: {actual_md5}")
关键注意点
- 优先用列表传参:彻底避免shell解析的编码问题和注入风险,是Python官方推荐的subprocess用法。
- 明确编码匹配:FFP文件和文件名使用的是Windows-1252编码,处理时需全程保持编码一致,不要混用工UTF-8和bytes。
- 避免强制text模式:不确定输出编码时,直接处理bytes,或通过
encoding参数指定对应编码,替代universal_newlines=True。
内容的提问来源于stack exchange,提问作者MonkeyMan47
相关产品推荐
相关产品推荐

