You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将DataFrame中含(keyword,score)对的列拆分为两列

问题

我有一个包含两列的DataFrame,一列是源文本,另一列是KeyBERT生成的(keyword, score)元组序列,示例如下:

[('white butterfly', 0.4587),
 ('pest horseradish', 0.4974),
 ('mature caterpillars', 0.6484)]

我希望将这列元组序列拆分为两列:一列存储所有关键词,另一列存储对应分数,目标格式如下:

df2data = [["The word horseradish is attested in English from the 1590s...",
            ['mispronunciation german', 'mareradish hypothesis', 'horse used'],
            [0.3715, 0.422, 0.4594]],
           ["Widely introduced by accident...",
            ['white butterfly', 'pest horseradish', 'mature caterpillars'],
            [0.4587, 0.4974, 0.6484]]]
df2 = pd.DataFrame(data = df2data, columns=['texts', 'words', 'scores'])

对应的表格形式:

wordsscorestexts
[mispronunciation german, mareradish hypothesi...][0.3715, 0.422, 0.4594]The word horseradish is attested in English from the 1590s...
[white butterfly, pest horseradish, mature cat...][0.4587, 0.4974, 0.6484]Widely introduced by accident...

我尝试过对words列做索引操作,但只能在序列内部遍历,没达到预期效果。

可复现示例代码:

import pandas as pd
from keybert import KeyBERT
kw_model = KeyBERT()

text = ["The word horseradish is attested in English from the 1590s...","Widely introduced by accident..."]
df = pd.DataFrame(text, columns=['texts'])
df['words'] = kw_model.extract_keywords(df['texts'], keyphrase_ngram_range=(1, 2), stop_words='english',
                              use_maxsum=True, nr_candidates=20, top_n=3)
解决方案

可以通过pandas的apply方法配合列表推导式,快速拆分元组序列为关键词列表和分数列表:

方法一:直接生成新列

# 先保留原始元组列的副本
df['temp_tuples'] = df['words'].copy()
# 提取关键词列表替换原words列
df['words'] = df['temp_tuples'].apply(lambda x: [item[0] for item in x])
# 提取分数列表生成新列
df['scores'] = df['temp_tuples'].apply(lambda x: [item[1] for item in x])
# 删除临时列
df.drop('temp_tuples', axis=1, inplace=True)
# 调整列顺序为目标格式
df = df[['texts', 'words', 'scores']]

方法二:用临时DataFrame合并(更简洁)

# 对每个元组序列生成包含words和scores的Series
temp_df = df['words'].apply(
    lambda x: pd.Series({
        'words': [item[0] for item in x],
        'scores': [item[1] for item in x]
    })
)
# 合并原文本列和临时DataFrame
df = pd.concat([df['texts'], temp_df], axis=1)

执行以上任意一种方法后,你的DataFrame就会变成目标格式:每行的words列是对应关键词的列表,scores列是匹配的分数列表。

内容的提问来源于stack exchange,提问作者SlowBear

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 20:42:07