You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python调用subprocess.run遇非UTF-8字符触发UnicodeDecodeError求助

问题分析

你的报错核心原因有两个:

  1. 设置了universal_newlines=True(等价于text=True),会强制subprocess将命令的stdout/stderr按系统默认编码(通常是UTF-8)解码为字符串,但文件名中的0xbb字节(对应Windows-1252编码的»)不符合UTF-8规则,触发了解码失败。
  2. 用bytes拼接命令字符串+shell=True的方式不仅不安全(存在shell注入风险),还容易因编码不匹配导致解析错误。
解决方案

优先采用列表传参+明确编码处理的方案,这是subprocess的最佳实践,能彻底规避编码和shell解析问题。

方案1:使用列表传参(推荐)

直接通过列表传递命令参数,subprocess会自动处理文件名的特殊字符,无需手动转义或拼接字符串:

import subprocess
from subprocess import PIPE

def test():
    # 先确定FFP文件的编码:从文件内容的`�`来看,是Windows-1252(cp1252)编码
    ffp_encoding = "cp1252"
    directory = "<redacted>"  # 用字符串形式,确保目录名编码匹配系统环境
    ffp_path = f"{directory}/<redacted>.ffp"
    
    # 用正确编码读取FFP文件,直接解析为字符串
    with open(ffp_path, 'r', encoding=ffp_encoding) as ffp_open:
        ffp_lines = ffp_open.readlines()
    
    for line in ffp_lines:
        line = line.strip()
        if not line.startswith(';') and ':' in line:
            filename_part, expected_md5 = line.split(':', 1)
            # 构造命令列表,自动处理文件名特殊字符
            cmd = [
                '/usr/bin/metaflac',
                '--show-md5sum',
                f"{directory}/{filename_part}"
            ]
            # 关闭text模式,直接获取bytes输出,或指定对应编码解码
            process = subprocess.run(
                cmd, 
                stdout=PIPE, 
                stderr=PIPE,
                encoding=ffp_encoding  # 直接指定编码,避免手动解码
            )
            actual_md5 = process.stdout.strip()
            print(f"预期指纹: {expected_md5.strip()}, 实际指纹: {actual_md5}")

方案2:保留shell=True但关闭text模式

如果必须使用shell命令拼接,需关闭universal_newlines=True,直接处理bytes输出:

import subprocess
from subprocess import PIPE

def test():
    directory = b'<redacted>'
    ffp_open = open(directory + b'<redacted>.ffp','rb')
    ffp_lines = ffp_open.readlines()
    
    for line in ffp_lines:
        if not line.startswith(b';') and b':' in line:
            txt = line.split(b':', 1)
            # 拼接bytes命令字符串(注意:若文件名含单引号会失效,需额外处理)
            ffp_cmd = b'/usr/bin/metaflac --show-md5sum \'' + directory + b'/' + txt[0].strip() + b'\''
            # 去掉universal_newlines=True,默认处理bytes
            process = subprocess.run(ffp_cmd, stdout=PIPE, stderr=PIPE, shell=True)
            actual_md5 = process.stdout.strip()
            expected_md5 = txt[1].strip()
            print(f"预期指纹: {expected_md5}, 实际指纹: {actual_md5}")
关键注意点
  • 优先用列表传参:彻底避免shell解析的编码问题和注入风险,是Python官方推荐的subprocess用法。
  • 明确编码匹配:FFP文件和文件名使用的是Windows-1252编码,处理时需全程保持编码一致,不要混用工UTF-8和bytes。
  • 避免强制text模式:不确定输出编码时,直接处理bytes,或通过encoding参数指定对应编码,替代universal_newlines=True。

内容的提问来源于stack exchange,提问作者MonkeyMan47

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 02:00:00