You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用facebook/m2m100_418M模型翻译长文本序列?

解决facebook/m2m100_418M翻译长文本时的截断问题

问题背景

使用facebook/m2m100_418M翻译维基百科上的Wiki定义长文本时,出现翻译内容截断的情况。对比原文和回译后的文本可以确认,这是模型的序列长度限制导致输入被截断。

解决方案

核心思路

要翻译更长的文本序列,最可靠的方式是分块翻译:把长文本拆分成符合模型最大序列长度的小块,逐一翻译后拼接结果。为避免上下文丢失,不能按固定字符数硬截断,而是要基于句子结尾分割,保证每个块都是完整的句子组合。

具体实现步骤

1. 安装依赖

先安装必要的库:

pip install transformers nltk

2. 完整代码实现

from transformers import M2M100Tokenizer, M2M100ForConditionalGeneration
import nltk
from nltk.tokenize import sent_tokenize

# 首次运行需要下载nltk的句子分割数据集
nltk.download('punkt')

# 加载预训练模型和分词器
model_name = "facebook/m2m100_418M"
tokenizer = M2M100Tokenizer.from_pretrained(model_name)
model = M2M100ForConditionalGeneration.from_pretrained(model_name)

# 获取模型支持的最大序列长度(m2m100_418M默认是1024)
MAX_SEQ_LENGTH = model.config.max_position_embeddings

def split_text_to_sentences(text):
    """把输入文本分割成独立句子列表"""
    return sent_tokenize(text)

def build_sentence_chunks(sentences, tokenizer, max_len):
    """将句子组合成不超过最大token长度的块,避免跨句子截断"""
    chunks = []
    current_chunk = []
    current_token_count = 0

    for sentence in sentences:
        # 计算当前句子的token数量
        sent_token_len = len(tokenizer.encode(sentence))
        # 预留特殊符号(<s>和</s>)的位置,所以要减2
        if current_token_count + sent_token_len > max_len - 2:
            # 当前块已满,存入列表并重置
            chunks.append(" ".join(current_chunk))
            current_chunk = [sentence]
            current_token_count = sent_token_len
        else:
            # 加入当前块,更新token计数
            current_chunk.append(sentence)
            current_token_count += sent_token_len
    # 把最后一个剩余的块加入列表
    if current_chunk:
        chunks.append(" ".join(current_chunk))
    return chunks

def translate_long_text(text, src_lang="eng", tgt_lang="fr"):
    """长文本翻译主函数,自动分块并拼接结果"""
    # 设置源语言和目标语言
    tokenizer.src_lang = src_lang
    forced_bos_token_id = tokenizer.get_lang_id(tgt_lang)

    # 分割句子并生成翻译块
    sentences = split_text_to_sentences(text)
    chunks = build_sentence_chunks(sentences, tokenizer, MAX_SEQ_LENGTH)

    # 逐块翻译
    translated_parts = []
    for chunk in chunks:
        inputs = tokenizer(chunk, return_tensors="pt")
        generated_tokens = model.generate(**inputs, forced_bos_token_id=forced_bos_token_id)
        translated_chunk = tokenizer.batch_decode(generated_tokens, skip_special_tokens=True)[0]
        translated_parts.append(translated_chunk)
    
    # 拼接所有翻译结果
    return " ".join(translated_parts)

# 测试示例:传入你的维基百科长文本
input_text = """A wiki is a form of hypertext publication on the internet which is collaboratively edited and managed by its audience directly through a web browser. A typical wiki contains multiple pages that can either be edited by the public or limited to use within an organization for maintaining its internal knowledge base.

Wikis are powered by wiki software, also known as wiki engines. Being a form of content management system, these differ from other web-based systems such as blog software or static site generators in that the content is created without any defined owner or leader. Wikis have little inherent structure, allowing one to emerge according to the needs of the users. Wiki engines usually allow content to be written using a lightweight markup language and sometimes edited with the help of a rich-text editor. There are dozens of different wiki engines in use, both standalone and part of other software, such as bug tracking systems. Some wiki engines are free and open-source, whereas others are proprietary. Some permit control over different functions (levels of access); for example, editing rights may permit changing, adding, or removing material. Others may permit access without enforcing access control. Further rules may be imposed to organize content. In addition to hosting user-authored content, wikis allow those users to interact, hold discussions, and collaborate."""

# 执行翻译并打印结果
final_translation = translate_long_text(input_text)
print(final_translation)

3. 关键细节说明

  • 句子分割:用nltk的sent_tokenize实现,能准确识别英文句子的结尾,保证每个块都是完整语义单元
  • 分块逻辑:通过计算句子的token长度,动态组合句子,确保每个块的总token数不超过模型上限,同时预留特殊符号的位置
  • 逐块翻译:沿用原模型的调用方式,只是把长文本拆成小块处理,最后拼接结果,不会丢失上下文信息

内容的提问来源于stack exchange,提问作者Naveen Reddy Marthala

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 11:24:49