You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

本地Python翻译HTML文件遇请求大小限制,求替代方案

解决HTML文件翻译的Payload大小限制问题及替代方案

先修复原Google Cloud Translate方案

Google Cloud Translate v2的单请求payload限制为200KB,你可以拆分HTML内容,仅提取并翻译文本节点,再重新组装回HTML结构,避免整段提交过大内容。示例代码:

from google.cloud import translate_v2
from bs4 import BeautifulSoup
import os

os.environ['GOOGLE_APPLICATION_CREDENTIALS'] = r"trans.json"
translate_client = translate_v2.Client()

with open("index.html", "r", encoding="utf-8") as f:
    html_content = f.read()

soup = BeautifulSoup(html_content, "html.parser")

# 遍历所有文本节点进行翻译
for text_node in soup.find_all(string=True):
    if text_node.strip():  # 跳过空白文本
        translated = translate_client.translate(
            text_node.strip(),
            target_language='hi',
            format_='text'
        )
        text_node.replace_with(translated['translatedText'])

# 保存翻译后的HTML
with open("translated_index.html", "w", encoding="utf-8") as f:
    f.write(str(soup))

替代方案1:DeepL Python库

DeepL的免费API支持更大的请求容量(免费版单请求最多500KB),且原生支持HTML格式翻译,无需手动拆分。

首先安装依赖:

pip install deepl

示例代码:

import deepl

# 初始化客户端(免费版用auth_key,付费版用auth_key或api_key)
translator = deepl.Translator("YOUR_DEEPL_AUTH_KEY")

with open("index.html", "r", encoding="utf-8") as f:
    html_content = f.read()

# 翻译HTML内容,指定目标语言为印地语(hi)
result = translator.translate_text(
    html_content,
    target_lang="HI",
    tag_handling="html"
)

with open("translated_index.html", "w", encoding="utf-8") as f:
    f.write(result.text)

替代方案2:OpenAI API

利用GPT系列模型的文本处理能力,提示模型仅翻译HTML中的文本内容,保留标签结构。

安装依赖:

pip install openai

示例代码:

from openai import OpenAI

client = OpenAI(api_key="YOUR_OPENAI_API_KEY")

with open("index.html", "r", encoding="utf-8") as f:
    html_content = f.read()

prompt = f"""请翻译以下HTML内容中的所有英文文本为印地语,严格保留HTML标签结构,不要修改任何标签:
{html_content}
"""

response = client.chat.completions.create(
    model="gpt-3.5-turbo",
    messages=[{"role": "user", "content": prompt}]
)

translated_html = response.choices[0].message.content

with open("translated_index.html", "w", encoding="utf-8") as f:
    f.write(translated_html)

替代方案3:基于国内翻译引擎的translate库

如果需要使用百度、有道等国内翻译引擎,可使用translate库,支持多引擎切换,需提前申请对应平台的API密钥。

安装依赖:

pip install translate

示例代码(以百度翻译为例):

from translate import Translator
from bs4 import BeautifulSoup

# 初始化百度翻译客户端,需填入你的APP ID和密钥
translator = Translator(
    to_lang="hi",
    provider="baidu",
    appid="YOUR_BAIDU_APPID",
    secret="YOUR_BAIDU_SECRET"
)

with open("index.html", "r", encoding="utf-8") as f:
    html_content = f.read()

# 拆分文本节点翻译后重组HTML
soup = BeautifulSoup(html_content, "html.parser")
for text_node in soup.find_all(string=True):
    if text_node.strip():
        translated_text = translator.translate(text_node.strip())
        text_node.replace_with(translated_text)

with open("translated_index.html", "w", encoding="utf-8") as f:
    f.write(str(soup))

内容的提问来源于stack exchange,提问作者visiwat303

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 19:10:33