使用Whisper AI转录波兰语时遭遇Unicode编码错误求助
问题概述
使用OpenAI Whisper工具转录波兰语音频文件时,因波兰语特殊字符(如ś、ż等)触发Unicode编码错误,但转录英语音频时同类型命令可正常运行。
使用的命令:whisper audio.wav --language Polish > transcriptPL.txt
(也曾尝试不指定波兰语;英语转录命令whisper audioEnglish.wav > transcriptEN.txt可正常执行)
错误信息:
C:\Users\agata\Desktop\Transcripts WP4> whisper audio.wav > transcriptPL.txt C:\Python312\Lib\site-packages\whisper\transcribe.py:115: UserWarning: FP16 is not supported on CPU; using FP32 instead warnings.warn("FP16 is not supported on CPU; using FP32 instead") Traceback (most recent call last): File "C:\Python312\Lib\site-packages\whisper\transcribe.py", line 478, in cli result = transcribe(model, audio_path, temperature=temperature, **args) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "C:\Python312\Lib\site-packages\whisper\transcribe.py", line 349, in transcribe print(make_safe(line)) File "C:\Python312\Lib\encodings\cp1252.py", line 19, in encode return codecs.charmap_encode(input,self.errors,encoding_table)[0] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ UnicodeEncodeError: 'charmap' codec can't encode characters in position 110-111: character maps to <undefined>
问题原因
Windows命令提示符(CMD)默认使用CP1252编码,该编码不支持波兰语的特殊字符。当Whisper输出包含这些字符的文本时,CMD的重定向功能无法处理,从而触发编码错误。英语字符均在CP1252编码范围内,因此不会出现问题。
解决方法
方法1:切换CMD编码为UTF-8
在执行Whisper命令前,先运行以下命令修改CMD的输出编码:chcp 65001
随后执行转录命令:whisper audio.wav --language Polish > transcriptPL.txt
方法2:使用Whisper自带的输出参数
通过Whisper内置的输出功能直接生成文件,绕过CMD的重定向限制:whisper audio.wav --language Polish --output_format txt --output_dir .
执行后,当前目录会生成audio.txt文件,完整包含支持Unicode字符的转录内容。
方法3:改用PowerShell执行命令
PowerShell默认支持UTF-8编码,直接在PowerShell中运行原命令即可:whisper audio.wav --language Polish > transcriptPL.txt
内容的提问来源于stack exchange,提问作者Agata T

