You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于PyTorch的问答任务数据集预处理报错排查

问题:HuggingFace问答任务预处理报错TypeError: 'list' object is not callable

背景信息

拥有如下格式的JSON数据集:

[{
    "answer": "...",
    "question": "...",
    "context": "..."
  },
 ...
]

目标是基于PyTorch用预训练BERT完成问答任务,参考HuggingFace文档实现预处理函数,并运行以下代码:

预处理函数

def preprocess_function(examples):
    questions = [q.strip() for q in examples["question"]]
    inputs = tokenizer(
        questions,
        examples["context"],
        max_length=384,
        truncation="only_second",
        return_offsets_mapping=True,
        padding="max_length",
    )

    offset_mapping = inputs.pop("offset_mapping")
    answers = examples["answers"]
    start_positions = []
    end_positions = []

    for i, offset in enumerate(offset_mapping):
        answer = answers[i]
        start_char = answer["answer_start"][0]
        end_char = answer["answer_start"][0] + len(answer["text"][0])
        sequence_ids = inputs.sequence_ids(i)

        # 定位上下文的起止位置
        idx = 0
        while sequence_ids[idx] != 1:
            idx += 1
        context_start = idx
        while sequence_ids[idx] == 1:
            idx += 1
        context_end = idx - 1

        # 若答案不在上下文内,标记为(0,0)
        if offset[context_start][0] > end_char or offset[context_end][1] < start_char:
            start_positions.append(0)
            end_positions.append(0)
        else:
            # 找到答案对应的token起止位置
            idx = context_start
            while idx <= context_end and offset[idx][0] <= start_char:
                idx += 1
            start_positions.append(idx - 1)

            idx = context_end
            while idx >= context_start and offset[idx][1] >= end_char:
                idx -= 1
            end_positions.append(idx + 1)

    inputs["start_positions"] = start_positions
    inputs["end_positions"] = end_positions
    return inputs

运行代码

from transformers import AutoTokenizer
from datasets import Dataset

dataset = Dataset.from_pandas(df) # df是包含answer、question、context三列的DataFrame

tokenizer = AutoTokenizer.from_pretrained("distilbert-base-uncased")
tokenized_ds = dataset.map(preprocess_function, batched=True)

报错信息

File "c:/project.py", line 234, in preprocess_function
    inputs = tokenizer(
TypeError: 'list' object is not callable

已确认数据集已转为Dataset对象,map函数使用正确,但错误仍存在。


错误分析

这个错误的核心原因是:代码中名为tokenizer的变量被赋值成了列表,而非预期的AutoTokenizer实例。当预处理函数尝试调用tokenizer()时,实际是在调用一个列表,因此触发"list不可调用"的错误。


解决方案

1. 排查变量名冲突

检查代码全局范围,确保没有其他地方将tokenizer变量重新赋值为列表(比如tokenizer = [...]这类代码),覆盖了之前通过AutoTokenizer.from_pretrained创建的实例。

2. 验证tokenizer实例类型

在实例化tokenizer后添加打印代码,确认其类型:

tokenizer = AutoTokenizer.from_pretrained("distilbert-base-uncased")
print(type(tokenizer))

正常输出应为类似<class 'transformers.models.distilbert.tokenization_distilbert.DistilBertTokenizer'>,如果输出是<class 'list'>,说明实例化过程存在问题或变量被覆盖。

3. 显式传入tokenizer到预处理函数

避免作用域导致的变量访问问题,使用functools.partial将tokenizer作为参数传入预处理函数:

from functools import partial

def preprocess_function(examples, tokenizer):
    questions = [q.strip() for q in examples["question"]]
    inputs = tokenizer(
        questions,
        examples["context"],
        max_length=384,
        truncation="only_second",
        return_offsets_mapping=True,
        padding="max_length",
    )
    # 剩余预处理逻辑保持不变...

# 调用map时传入tokenizer
tokenized_ds = dataset.map(partial(preprocess_function, tokenizer=tokenizer), batched=True)

4. 适配数据集格式到Squad规范

原预处理函数基于Squad格式的answers字段(包含answer_start和text的字典)编写,但你的原始数据集只有纯文本的answer字段,需要先转换格式:

def convert_to_squad_format(example):
    answer_text = example["answer"]
    # 在上下文中查找答案的起始位置
    start_idx = example["context"].find(answer_text)
    # 处理答案不在上下文中的情况,可根据需求调整
    if start_idx == -1:
        return {"answers": {"answer_start": [0], "text": [""]}}
    return {"answers": {"answer_start": [start_idx], "text": [answer_text]}}

# 转换数据集格式
dataset = dataset.map(convert_to_squad_format)

内容的提问来源于stack exchange,提问作者evader110

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 12:26:20