如何在cuDF中高效将不同类型列配对成元组列表
如何在cuDF中保持设备端数据的前提下合并字符串列与列表列
问题背景
我有如下cuDF DataFrame:
Context Questions 0 Architecturally, the school has a Catholic cha... [To whom did the Virgin Mary allegedly appear ... 1 As at most other universities, Notre Dame's st... [When did the Scholastic Magazine of Notre dam... 2 The university is the major seat of the Congre... [Where is the headquarters of the Congregation... 3 The College of Engineering was established in ... [How many BS level degrees are offered in the ... 4 All of Notre Dame's undergraduate students are... [What entity provides help with the management...
其中Context列是单个字符串,Questions列是字符串列表。我需要生成一个新列,内容为类似[(Context, question_i)]的配对列表。
针对SQuaD-v.1数据集,我曾用以下代码处理Questions列:
data = cudf.read_csv(DATA_PATH) pattern = '([^"]+\?)' data["Questions"] = data['QuestionAnswerSets'].str.replace('Question" -> "', '').str.findall(pattern)
遇到的问题:
- 不想调用列表构造函数,这会导致数据从设备端(device)转移到主机端(host),失去GPU加速优势。
- 尝试自定义UDF时报错,因为UDF无法兼容当前的数据类型组合:
def zip_context_question_pairs(row): return row['Context'], row['Questions'] df = df.apply_rows(zip_context_question_pairs, incols=['Context', 'Questions'], outcols={'Context_QuestionPairs': 'object'}, kwargs={})
可复现代码:
df = cudf.DataFrame({ 'context': 'Architecturally, the school has a Catholic character.', 'question': [['To whom did the Virgin Mary allegedly appear?', "another question"]], }) context = df["context"][0] questions = df["question"][0] desired_result = [] # 希望用cuDF方法替代该循环,避免数据转移到主机端 for question in questions : desired_result.append((question, context)) print(desired_result)
解决方案
核心思路:用explode+groupby实现设备端操作
完全基于cuDF原生API,无需将数据转移到主机端,步骤如下:
- 展开列表列:使用
explode将Questions列的每个列表元素拆分为单独行,同时保留对应行的Context值。 - 分组聚合配对:按原行索引分组,将每组的
(Context, Question)对重新组合为列表。
代码实现:
# 1. 展开question列,每个问题对应一行,保留原索引 exploded_df = df.explode('question') # 2. 按原索引分组,聚合生成(Context, question)配对列表 result_df = exploded_df.groupby(exploded_df.index).apply( lambda g: list(zip(g['context'], g['question'])), meta=('Context_QuestionPairs', 'list') ) # 3. 将结果合并回原DataFrame df = df.join(result_df)
执行后,df['Context_QuestionPairs']即为所需的配对列表列,整个过程完全在GPU设备端完成,不会触发数据到CPU的转移。
方案优势
- 完全使用cuDF原生操作,避免自定义UDF的类型兼容问题。
- 全程GPU执行,保持数据在设备端,维持计算效率。
- 输出的列表列可直接用于后续的GPU加速操作。
内容的提问来源于stack exchange,提问作者JOKKINATOR
相关产品推荐
相关产品推荐

