Python 3调用FFmpeg子进程获取stderr时出现字符解码错误
解决subprocess调用FFmpeg处理日文字幕时的编码解码错误
问题根源
你遇到的'charmap' codec can't decode byte错误,本质是编码不匹配:
- 设置
universal_newlines=True后,subprocess会使用系统默认编码(Windows下通常是GBK/GB2312)解码FFmpeg的stderr输出 - 但FFmpeg的标准错误输出默认采用UTF-8编码,日文字符的字节流用系统默认编码解码就会出现无法识别的字符,触发解码失败
解决方案
方案1:直接指定编码(Python 3.7+)
Python 3.7及以上版本的subprocess.Popen支持encoding参数,直接指定为utf-8即可:
command = [ffmpeg, "-y", "-i", fileSelected, "-acodec", "pcm_s16le", "-vn", "-t", "3", "-f", "null", "-"] print(command) # 替换universal_newlines=True为encoding='utf-8',同时修正Jstdin的拼写错误 proc = subprocess.Popen(command, stderr=subprocess.PIPE, stdin=subprocess.PIPE, encoding='utf-8', startupinfo=startupinfo) stream = "" for line in proc.stderr: try: print("line", line) if "Stream #" in line: estream = line.split("#",1)[1] estream = estream.split(" (",1)[0] print("estream", estream) stream += estream + "\n" except Exception as error: # 修正exception为大写Exception print("处理错误:", error)
方案2:手动解码字节流(兼容Python 3.6及以下)
如果你的Python版本低于3.7,不支持encoding参数,可以直接读取字节流,再手动用UTF-8解码:
command = [ffmpeg, "-y", "-i", fileSelected, "-acodec", "pcm_s16le", "-vn", "-t", "3", "-f", "null", "-"] print(command) # 去掉universal_newlines=True,直接读取字节 proc = subprocess.Popen(command, stderr=subprocess.PIPE, stdin=subprocess.PIPE, startupinfo=startupinfo) stream = "" for line_bytes in proc.stderr: try: # 用UTF-8解码字节流,strip()去除换行符 line = line_bytes.decode('utf-8').strip() print("line", line) if "Stream #" in line: estream = line.split("#",1)[1] estream = estream.split(" (",1)[0] print("estream", estream) stream += estream + "\n" except UnicodeDecodeError as error: print("解码错误:", error) # 可选:用replace模式忽略无法解码的字符,避免程序中断 line = line_bytes.decode('utf-8', errors='replace').strip() # 后续可继续处理line内容
额外注意事项
- 修正原代码中的拼写错误:
Jstdin应为stdin,exception应为大写Exception - 如果遇到极少数FFmpeg输出编码不是UTF-8的情况,可以尝试
shift-jis编码(日文常见编码)进行解码
内容的提问来源于stack exchange,提问作者jdauthre
相关产品推荐
相关产品推荐

