You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Hugging Face数据集循环添加summary键值对无效的解决方法咨询

问题描述

我有一个Hugging Face数据集,示例格式如下:

test = [{'doc': document1, 'id': id1}, {'doc': document2, 'id': id2}.......]

我尝试为每条数据生成LexRank摘要,并添加为新的summary键值对,但执行以下代码后,数据集无任何变化且未报错:

for i in test:
   segments = get_segmented_text(i['doc'])
   expected_length = round(len(segments) / median_compression_ratio)
   most_central_indices = compute_lexrank_sentences(model, segments, device, expected_length)
   summary = [segments[idx] for idx in sorted(most_central_indices)]
   i['summary'] = '\n'.join(summary)

期望得到的数据集格式:

test = [{'doc': document1, 'id': id1, 'summary': summary1}, {'doc': document2, 'id': id2, 'summary': summary2}.......]

补充说明:将数据集转为DataFrame后,以下代码可以正常运行,但我希望直接在Hugging Face数据集上实现该操作:

for i, row in df.iterrows():
    segments = get_segmented_text(row['doc'])
    expected_length = round(len(segments) / median_compression_ratio)
    most_central_indices = compute_lexrank_sentences(model, segments, device, expected_length)

    summary = [segments[idx] for idx in sorted(most_central_indices)]
    df.at[i, 'summary'] = '\n'.join(summary)
解决方案

问题根源是Hugging Face的Dataset对象是**不可变(immutable)**的:你遍历数据集时拿到的是数据项的拷贝,而非原数据集的引用,因此修改拷贝不会影响原数据集。以下是两种直接操作Hugging Face数据集的有效方法:

方法1:使用Dataset.map()(官方推荐)

这是处理Hugging Face数据集的标准方式,效率高且符合最佳实践。把生成摘要的逻辑封装成函数,通过map方法批量为每条数据添加summary字段:

def add_summary(example):
    segments = get_segmented_text(example['doc'])
    expected_length = round(len(segments) / median_compression_ratio)
    most_central_indices = compute_lexrank_sentences(model, segments, device, expected_length)
    summary = [segments[idx] for idx in sorted(most_central_indices)]
    example['summary'] = '\n'.join(summary)
    return example

# 应用到数据集
test_dataset = test_dataset.map(add_summary)

如果数据集规模较大,可添加batched=True参数开启批量处理(需调整函数适配批量输入),进一步提升处理速度。

方法2:转为可变列表后修改

如果偏好循环方式,可先将Dataset转为普通Python列表(可变类型),修改后再转回Dataset:

# 将Dataset转为Python列表
test_list = test_dataset.to_list()

# 循环修改每个字典项
for item in test_list:
    segments = get_segmented_text(item['doc'])
    expected_length = round(len(segments) / median_compression_ratio)
    most_central_indices = compute_lexrank_sentences(model, segments, device, expected_length)
    summary = [segments[idx] for idx in sorted(most_central_indices)]
    item['summary'] = '\n'.join(summary)

# 转回Hugging Face Dataset
from datasets import Dataset
test_dataset = Dataset.from_list(test_list)

原代码失效的原因

Hugging Face Dataset迭代时返回的是数据项的独立拷贝,而非原数据集的引用;而DataFrame的iterrows()返回的是可修改的行视图,因此修改操作能直接生效。

内容的提问来源于stack exchange,提问作者Praveen Bushipaka

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 03:25:19