基于PyTorch的问答任务数据集预处理报错排查
问题:HuggingFace问答任务预处理报错TypeError: 'list' object is not callable
背景信息
拥有如下格式的JSON数据集:
[{ "answer": "...", "question": "...", "context": "..." }, ... ]
目标是基于PyTorch用预训练BERT完成问答任务,参考HuggingFace文档实现预处理函数,并运行以下代码:
预处理函数
def preprocess_function(examples): questions = [q.strip() for q in examples["question"]] inputs = tokenizer( questions, examples["context"], max_length=384, truncation="only_second", return_offsets_mapping=True, padding="max_length", ) offset_mapping = inputs.pop("offset_mapping") answers = examples["answers"] start_positions = [] end_positions = [] for i, offset in enumerate(offset_mapping): answer = answers[i] start_char = answer["answer_start"][0] end_char = answer["answer_start"][0] + len(answer["text"][0]) sequence_ids = inputs.sequence_ids(i) # 定位上下文的起止位置 idx = 0 while sequence_ids[idx] != 1: idx += 1 context_start = idx while sequence_ids[idx] == 1: idx += 1 context_end = idx - 1 # 若答案不在上下文内,标记为(0,0) if offset[context_start][0] > end_char or offset[context_end][1] < start_char: start_positions.append(0) end_positions.append(0) else: # 找到答案对应的token起止位置 idx = context_start while idx <= context_end and offset[idx][0] <= start_char: idx += 1 start_positions.append(idx - 1) idx = context_end while idx >= context_start and offset[idx][1] >= end_char: idx -= 1 end_positions.append(idx + 1) inputs["start_positions"] = start_positions inputs["end_positions"] = end_positions return inputs
运行代码
from transformers import AutoTokenizer from datasets import Dataset dataset = Dataset.from_pandas(df) # df是包含answer、question、context三列的DataFrame tokenizer = AutoTokenizer.from_pretrained("distilbert-base-uncased") tokenized_ds = dataset.map(preprocess_function, batched=True)
报错信息
File "c:/project.py", line 234, in preprocess_function inputs = tokenizer( TypeError: 'list' object is not callable
已确认数据集已转为Dataset对象,map函数使用正确,但错误仍存在。
错误分析
这个错误的核心原因是:代码中名为tokenizer的变量被赋值成了列表,而非预期的AutoTokenizer实例。当预处理函数尝试调用tokenizer()时,实际是在调用一个列表,因此触发"list不可调用"的错误。
解决方案
1. 排查变量名冲突
检查代码全局范围,确保没有其他地方将tokenizer变量重新赋值为列表(比如tokenizer = [...]这类代码),覆盖了之前通过AutoTokenizer.from_pretrained创建的实例。
2. 验证tokenizer实例类型
在实例化tokenizer后添加打印代码,确认其类型:
tokenizer = AutoTokenizer.from_pretrained("distilbert-base-uncased") print(type(tokenizer))
正常输出应为类似<class 'transformers.models.distilbert.tokenization_distilbert.DistilBertTokenizer'>,如果输出是<class 'list'>,说明实例化过程存在问题或变量被覆盖。
3. 显式传入tokenizer到预处理函数
避免作用域导致的变量访问问题,使用functools.partial将tokenizer作为参数传入预处理函数:
from functools import partial def preprocess_function(examples, tokenizer): questions = [q.strip() for q in examples["question"]] inputs = tokenizer( questions, examples["context"], max_length=384, truncation="only_second", return_offsets_mapping=True, padding="max_length", ) # 剩余预处理逻辑保持不变... # 调用map时传入tokenizer tokenized_ds = dataset.map(partial(preprocess_function, tokenizer=tokenizer), batched=True)
4. 适配数据集格式到Squad规范
原预处理函数基于Squad格式的answers字段(包含answer_start和text的字典)编写,但你的原始数据集只有纯文本的answer字段,需要先转换格式:
def convert_to_squad_format(example): answer_text = example["answer"] # 在上下文中查找答案的起始位置 start_idx = example["context"].find(answer_text) # 处理答案不在上下文中的情况,可根据需求调整 if start_idx == -1: return {"answers": {"answer_start": [0], "text": [""]}} return {"answers": {"answer_start": [start_idx], "text": [answer_text]}} # 转换数据集格式 dataset = dataset.map(convert_to_squad_format)
内容的提问来源于stack exchange,提问作者evader110
相关产品推荐
相关产品推荐

