You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解决Hugging Face翻译预处理中的TypeError: int()参数为NoneType错误

问题:Hugging Face翻译预处理代码运行报错TypeError

我运行以下取自Hugging Face官网的翻译预处理代码:

EN_AR = load_dataset("iwslt2017", "iwslt2017-ar-en", split="train").select(range(2000))

def extract_languages(examples):
  inputs = [ex["ar"] for ex in examples['translation']]
  target = [ex["en"] for ex in examples['translation']]
  return {"inputs":inputs,"targets":target}

EN_AR = EN_AR.map(extract_languages,batched=True, remove_columns=["translation"])

from transformers import AutoTokenizer, MBart50TokenizerFast

model_name = "facebook/mbart-large-50"
tokenizer = AutoTokenizer.from_pretrained(model_name)
maxL = 128
def preprocess_func(examples):
  model_inputs = tokenizer(examples["inputs"],max_length=maxL,truncation=True)

  with tokenizer.as_target_tokenizer():
    labels = tokenizer(examples["targets"],max_length=maxL,truncation=True)

  model_inputs["labels"]= labels["input_ids"]
  return model_name

tokenized_datasets = EN_AR.map(preprocess_func, batched = True, remove_columns=["inputs","targets"])

收到如下错误:

TypeError                                 Traceback (most recent call last)
 in <cell line: 15>()
13   return model_name
14
---> 15 tokenized_datasets = EN_AR.map(preprocess_func, batched = True, remove_columns=["inputs","targets"])

10 frames
/usr/local/lib/python3.10/dist-packages/transformers/tokenization_utils_fast.py in convert_ids_to_tokens(self, ids, skip_special_tokens)
387         tokens = []
388         for index in ids:
---> 389             index = int(index)
390             if skip_special_tokens and index in self.all_special_ids:
391                 continue

TypeError: int() argument must be a string, a bytes-like object or a real number, not 'NoneType'

请求排查并解决该错误。


错误原因与修复方案

错误根源

preprocess_func函数最后错误地返回了model_name(字符串),而不是处理好的model_inputs字典。Dataset的map方法要求每个批次处理后返回包含特征的字典,返回字符串会导致后续处理逻辑混乱,触发类型转换错误。

修正后的代码

将preprocess_func的返回值改为model_inputs即可:

EN_AR = load_dataset("iwslt2017", "iwslt2017-ar-en", split="train").select(range(2000))

def extract_languages(examples):
  inputs = [ex["ar"] for ex in examples['translation']]
  target = [ex["en"] for ex in examples['translation']]
  return {"inputs":inputs,"targets":target}

EN_AR = EN_AR.map(extract_languages,batched=True, remove_columns=["translation"])

from transformers import AutoTokenizer, MBart50TokenizerFast

model_name = "facebook/mbart-large-50"
tokenizer = AutoTokenizer.from_pretrained(model_name)
maxL = 128
def preprocess_func(examples):
  model_inputs = tokenizer(examples["inputs"],max_length=maxL,truncation=True)

  with tokenizer.as_target_tokenizer():
    labels = tokenizer(examples["targets"],max_length=maxL,truncation=True)

  model_inputs["labels"]= labels["input_ids"]
  # 修正:返回处理后的特征字典,而非model_name字符串
  return model_inputs

tokenized_datasets = EN_AR.map(preprocess_func, batched = True, remove_columns=["inputs","targets"])

额外优化建议

使用MBart50TokenizerFast时,建议显式指定源语言和目标语言的tokenizer设置,避免潜在歧义:

# 替换tokenizer初始化代码
tokenizer = MBart50TokenizerFast.from_pretrained(model_name, src_lang="ar_AR", tgt_lang="en_XX")

内容的提问来源于stack exchange,提问作者shahad

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 01:02:15