You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python3.9下Huggingface datasets的load_dataset报KeyError问题求助

解决Python 3.9环境下加载Huggingface数据集的KeyError问题

问题背景

在Ubuntu 22系统的Python 3.9环境中,执行以下代码加载SQuAD v2数据集时出现KeyError: 'length'错误,但Python 3.10环境可正常运行:

from datasets import load_dataset
dataset_squad_v2 = load_dataset("squad_v2")

报错详情:

lib/python3.9/site-packages/datasets/features/features.py in generate_from_dict(obj)
   1282 
   1283     if class_type == Sequence:
-> 1284         return Sequence(feature=generate_from_dict(obj["feature"]), length=obj["length"])
   1285 
   1286     field_names = {f.name for f in fields(class_type)}

KeyError: 'length'

已尝试强制重新下载数据集、更新datasets库,问题仍未解决,且加载其他Huggingface数据集时也存在类似问题。

解决方案

1. 降级datasets库到兼容Python 3.9的版本

高版本datasets(如>=2.10.0)在Python 3.9下存在特征解析的兼容性问题,可安装更早的稳定版本:

pip install datasets==2.9.0

若仍有问题,可尝试更低版本如2.8.0。

2. 降级pyarrow依赖

datasets依赖pyarrow处理数据序列化,高版本pyarrow与Python 3.9不兼容会导致该错误,安装兼容版本:

pip install pyarrow==11.0.0

安装完成后重新运行加载代码。

3. 手动定义数据集特征(备选方案)

如果上述方法无效,可手动指定SQuAD v2的特征结构,绕过自动推断逻辑:

from datasets import load_dataset, Features, Value, Sequence

# 手动定义SQuAD v2的特征结构
squad_v2_features = Features({
    "id": Value("string"),
    "title": Value("string"),
    "context": Value("string"),
    "question": Value("string"),
    "answers": Sequence({
        "text": Value("string"),
        "answer_start": Value("int32")
    })
})

# 加载数据集时指定自定义特征
dataset_squad_v2 = load_dataset("squad_v2", features=squad_v2_features)

4. 彻底清理数据集缓存

缓存的数据集元数据可能损坏,手动删除缓存目录后重新加载:

rm -rf ~/.cache/huggingface/datasets/squad_v2

内容的提问来源于stack exchange,提问作者Jan Spörer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 04:43:20