You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Google Translate API处理UTF16-BE文件时遇UnicodeEncodeError求助

解决谷歌翻译API写入UTF-16BE文件时的UnicodeEncodeError问题

问题还原

你在用Google Translate API处理UTF-16BE编码的文本文件时,大部分行都能正常处理,但偶尔会弹出UnicodeEncodeError,提示'charmap' codec can't encode character '\u0259',而且只有写入文件时报错,命令行打印完全正常。你的核心代码片段大致是这样的:

import json
import urllib
from urllib.request import Request, urlopen
import urllib.parse

def googletranslate(sourceLang, targetLang, sourceText):
    url = "https://translate.googleapis.com/translate_a/single?client=gtx&sl=" + sourceLang + "&tl=" + targetLang + "&dt=t&q=" + urllib.parse.quote_plus(sourceText)
    urld = Request(url, headers={'User-Agent': 'Mozilla/5.0'})
    jsonfile = urlopen(urld).read()
    h = json.loads(jsonfile)
    return h[0][0][0]

input = [line.rstrip('\n') for line in open('input.txt', 'r', encoding="utf_16_be")]
output = open('output.txt', 'w', encoding="utf_16_be")

# 假设offset和size是已定义的变量
for y in range(offset,offset+size):
    text = input[y]
    text = googletranslate('auto', '<desired language>', text)
    text.encode('utf_16_be')  # 这行代码没有实际作用
    print("T: " + text)
    output.write(text + '\n')

错误信息如下:

T: <translated text>
Traceback (most recent call last):
  File "C:\PATH\TO\translate.py", line 124, in <module>
    output.write(text + '\n')
  File "C:\PATH\TO\AppData\Local\Programs\Python\Python36-32\lib\encodings\cp1252.py", line 19, in encode
    return codecs.charmap_encode(input,self.errors,encoding_table)[0]
UnicodeEncodeError: 'charmap' codec can't encode character '\u0259' in position 22: character maps to <undefined>

错误原因分析

  1. 无效的编码操作:你代码里的text.encode('utf_16_be')只是计算了字节串,但没有赋值回text变量,所以这行代码完全没用,text依然是Unicode字符串。
  2. 文件编码设置被干扰:虽然你打开输出文件时指定了encoding="utf_16_be",但错误信息显示实际使用的是cp1252编码(Windows系统的默认ANSI编码)。这大概率是Windows环境下文本模式的编码优先级问题,或者Python环境变量的默认编码干扰了open函数的encoding参数。
  3. UTF-16BE完全能容纳目标字符:别担心,\u0259(中央元音ə)属于Unicode字符集,UTF-16BE完全可以表示它,问题不是编码不够用,而是写入时的编码逻辑出了问题。

解决方案

最可靠的方法是用二进制模式打开输出文件,手动将字符串编码为UTF-16BE字节串后写入,绕开文本模式下的编码自动处理:

  1. 修改输出文件的打开方式为二进制模式:
output = open('output.txt', 'wb')  # 'wb'表示二进制写入模式
  1. 在写入时,主动将字符串(包括换行符)编码为UTF-16BE:
for y in range(offset,offset+size):
    text = input[y]
    text = googletranslate('auto', '<desired language>', text)
    print("T: " + text)
    # 先拼接字符串,再编码为UTF-16BE字节串后写入
    output.write( (text + '\n').encode('utf_16_be') )
  1. 推荐用with语句自动管理文件(避免忘记关闭导致的问题):
# 读取输入文件
with open('input.txt', 'r', encoding="utf_16_be") as f:
    input_lines = [line.rstrip('\n') for line in f]

# 写入输出文件
with open('output.txt', 'wb') as output:
    for y in range(offset, offset+size):
        text = input_lines[y]
        text = googletranslate('auto', '<desired language>', text)
        print("T: " + text)
        output.write( (text + '\n').encode('utf_16_be') )

额外说明

  • 为什么命令行打印正常?因为你的命令行终端编码(比如UTF-8)支持\u0259这个字符,所以能正常显示,但文件写入时因为编码设置错误才报错。
  • 如果坚持用文本模式,也可以尝试在打开文件时加上errors='backslashreplace'或errors='replace'来处理无法编码的字符,但这会丢失或替换字符,不如二进制模式可靠:
output = open('output.txt', 'w', encoding="utf_16_be", errors='backslashreplace')

内容的提问来源于stack exchange,提问作者THO games

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 09:22:14