You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在cuDF中高效将不同类型列配对成元组列表

如何在cuDF中保持设备端数据的前提下合并字符串列与列表列

问题背景

我有如下cuDF DataFrame:

Context                                          Questions
0      Architecturally, the school has a Catholic cha...  [To whom did the Virgin Mary allegedly appear ...
1      As at most other universities, Notre Dame's st...  [When did the Scholastic Magazine of Notre dam...
2      The university is the major seat of the Congre...  [Where is the headquarters of the Congregation...
3      The College of Engineering was established in ...  [How many BS level degrees are offered in the ...
4      All of Notre Dame's undergraduate students are...  [What entity provides help with the management...

其中Context列是单个字符串,Questions列是字符串列表。我需要生成一个新列,内容为类似[(Context, question_i)]的配对列表。

针对SQuaD-v.1数据集,我曾用以下代码处理Questions列:

data = cudf.read_csv(DATA_PATH)
pattern = '([^"]+\?)'

data["Questions"] = data['QuestionAnswerSets'].str.replace('Question" -> "', '').str.findall(pattern)

遇到的问题:

  • 不想调用列表构造函数,这会导致数据从设备端(device)转移到主机端(host),失去GPU加速优势。
  • 尝试自定义UDF时报错,因为UDF无法兼容当前的数据类型组合:
def zip_context_question_pairs(row):
    return row['Context'], row['Questions']

df = df.apply_rows(zip_context_question_pairs,
                   incols=['Context', 'Questions'],
                   outcols={'Context_QuestionPairs': 'object'},
                   kwargs={})

可复现代码:

df = cudf.DataFrame({
    'context': 'Architecturally, the school has a Catholic character.',
    'question': [['To whom did the Virgin Mary allegedly appear?', "another question"]],
    })

context = df["context"][0]
questions = df["question"][0]

desired_result = []

# 希望用cuDF方法替代该循环,避免数据转移到主机端
for question in questions :
    desired_result.append((question, context))

print(desired_result)

解决方案

核心思路:用explode+groupby实现设备端操作

完全基于cuDF原生API,无需将数据转移到主机端,步骤如下:

  1. 展开列表列:使用explode将Questions列的每个列表元素拆分为单独行,同时保留对应行的Context值。
  2. 分组聚合配对:按原行索引分组,将每组的(Context, Question)对重新组合为列表。

代码实现:

# 1. 展开question列,每个问题对应一行,保留原索引
exploded_df = df.explode('question')

# 2. 按原索引分组,聚合生成(Context, question)配对列表
result_df = exploded_df.groupby(exploded_df.index).apply(
    lambda g: list(zip(g['context'], g['question'])),
    meta=('Context_QuestionPairs', 'list')
)

# 3. 将结果合并回原DataFrame
df = df.join(result_df)

执行后,df['Context_QuestionPairs']即为所需的配对列表列,整个过程完全在GPU设备端完成,不会触发数据到CPU的转移。

方案优势

  • 完全使用cuDF原生操作,避免自定义UDF的类型兼容问题。
  • 全程GPU执行,保持数据在设备端,维持计算效率。
  • 输出的列表列可直接用于后续的GPU加速操作。

内容的提问来源于stack exchange,提问作者JOKKINATOR

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 21:53:14