使用Hugging Face Dataset创建自定义数据集时遇TypeError问题求助
解决Hugging Face Dataset创建时的TypeError问题
错误原因
你遇到的TypeError: expected bytes, int found,核心问题是**train_pairs的结构不符合Dataset.from_dict的要求**。from_dict需要接收「字段名到对应样本列表」的映射(比如{'sentence1': [句1, 句2], 'label': [0, 1]}),而非单个样本的字典。即便把label改成整数,只要输入结构错误,就会触发类型解析异常。
修复步骤
1. 调整train_pairs的结构
确保每个字段对应包含所有样本的列表,示例如下:
train_pairs = { 'sentence1': [ "that 's far too tragic to merit such superficial treatment ", "another example sentence" ], 'sentence2': [ "that 's far too tragic to merit such superficial treatment ", "another paired sentence" ], 'label': [0, 1], # 用整数对应ClassLabel的索引 'idx': [5, 6] }
如果你的数据是单个样本组成的列表(比如[样本1字典, 样本2字典]),改用Dataset.from_list创建更直观:
train_samples = [ { 'sentence1': "that 's far too tragic to merit such superficial treatment ", 'sentence2': "that 's far too tragic to merit such superficial treatment ", 'label': 0, 'idx': 5 }, # 添加更多样本... ] custom_dataset = Dataset.from_list(train_samples)
2. 正确指定Features(两种方式)
方式一:创建Dataset时直接指定features
from datasets import Dataset, ClassLabel, Value features = { "sentence1": Value("string"), "sentence2": Value("string"), "label": ClassLabel(names=["not_equivalent", "equivalent"]), "idx": Value("int32"), } # 基于结构正确的train_pairs创建 custom_dataset = Dataset.from_dict(train_pairs, features=features) # 或者基于样本列表创建 custom_dataset = Dataset.from_list(train_samples, features=features)
方式二:创建后再执行cast(需确保Dataset结构正确)
如果已经创建了Dataset,可在结构验证无误后执行cast:
# 先创建结构正确的Dataset custom_dataset = Dataset.from_dict(train_pairs) # 再转换为目标features custom_dataset = custom_dataset.cast(features)
3. 验证结果
完成上述操作后,打印custom_dataset即可得到预期结构,示例输出:
Dataset({ features: ['sentence1', 'sentence2', 'label', 'idx'], num_rows: 2 })
查看单个样本时,label会自动关联ClassLabel的名称,比如执行custom_dataset.features['label'].int2str(0)会返回"not_equivalent"。
内容的提问来源于stack exchange,提问作者user269867
相关产品推荐
相关产品推荐

