使用googletrans将HTML翻译为印地语时遇两类错误,求修复方案
问题分析与修复方案
第一类错误:TypeError (JSON对象为NoneType)
错误原因
googletrans官方版本的API接口已失效,谷歌翻译返回格式变更后,库无法正确解析数据,返回None触发JSON解析错误;频繁请求也可能被谷歌临时限制,加剧该问题。
修复方案
- 更换稳定版本或替代库
卸载现有googletrans,安装修复过接口的测试版本:
或改用更稳定的pip uninstall googletrans pip install googletrans==4.0.0-rc1deep-translator库:pip install deep-translator - 增加请求延迟
在翻译循环中添加延迟,降低请求频率避免被限制:import time # 翻译操作后添加 time.sleep(1) - 精准捕获异常
避免使用裸except,只捕获请求和解析相关的特定异常,防止掩盖其他问题。
第二类错误:UnicodeEncodeError (控制台编码问题)
错误原因
Windows默认控制台编码为cp1252,而待打印文本中包含该编码不支持的Unicode字符(比如→,对应\u2192),导致打印时编码失败。该问题与文件读写编码无关,仅为控制台输出限制。
修复方案
- 强制控制台输出为UTF-8
在代码开头添加以下代码,修改标准输出编码:import sys sys.stdout.reconfigure(encoding='utf-8') - 处理打印时的编码异常
打印时忽略或替换无法编码的字符:print("Translation failed for element: ", element.encode('utf-8', errors='replace').decode('utf-8')) - 将错误写入日志文件
替代控制台打印,把错误信息写入日志文件,规避控制台编码限制:with open('translation_errors.log', 'a', encoding='utf-8') as log_f: log_f.write(f"Translation failed for element: {element}\n")
修改后的完整代码示例
以使用googletrans==4.0.0-rc1和修复控制台编码为例:
import os import sys import time from bs4 import BeautifulSoup from googletrans import Translator # 修复控制台输出编码问题 sys.stdout.reconfigure(encoding='utf-8') # 设置输入输出目录 input_dir = r"C:\My Web Sites\CC\www.classcentral.com\subject" output_dir = r"C:\My Web Sites\CC\www.classcentral.com\translated\subject" # 创建输出目录 if not os.path.exists(output_dir): os.makedirs(output_dir) # 初始化翻译器 translator = Translator(service_urls=['translate.google.com']) # 遍历HTML文件 for filename in os.listdir(input_dir): if filename.endswith('.html'): file_path = os.path.join(input_dir, filename) # 用utf-8读取文件更稳妥,避免latin-1丢失字符 with open(file_path, 'r', encoding='utf-8') as f: soup = BeautifulSoup(f, 'html.parser') # 翻译文本内容 for element in soup.find_all(text=True): if element.strip(): try: translated_text = translator.translate(element.string, dest='hi').text element.string.replace_with(translated_text) time.sleep(1) # 增加延迟避免被谷歌限制 except Exception as e: # 同时打印并记录错误日志 error_msg = f"文件[{filename}]中元素翻译失败: {element}\n错误详情: {str(e)}\n" print(error_msg) with open('translation_errors.log', 'a', encoding='utf-8') as log_f: log_f.write(error_msg) # 用utf-8写入翻译后的文件 output_file = os.path.join(output_dir, filename) with open(output_file, 'w', encoding='utf-8') as f: f.write(str(soup)) print(f"已完成翻译: {filename} -> {output_file}")
内容的提问来源于stack exchange,提问作者CartoonSavage32
相关产品推荐
相关产品推荐

