如何将DataFrame中含(keyword,score)对的列拆分为两列
问题
我有一个包含两列的DataFrame,一列是源文本,另一列是KeyBERT生成的(keyword, score)元组序列,示例如下:
[('white butterfly', 0.4587), ('pest horseradish', 0.4974), ('mature caterpillars', 0.6484)]
我希望将这列元组序列拆分为两列:一列存储所有关键词,另一列存储对应分数,目标格式如下:
df2data = [["The word horseradish is attested in English from the 1590s...", ['mispronunciation german', 'mareradish hypothesis', 'horse used'], [0.3715, 0.422, 0.4594]], ["Widely introduced by accident...", ['white butterfly', 'pest horseradish', 'mature caterpillars'], [0.4587, 0.4974, 0.6484]]] df2 = pd.DataFrame(data = df2data, columns=['texts', 'words', 'scores'])
对应的表格形式:
| words | scores | texts |
|---|---|---|
| [mispronunciation german, mareradish hypothesi...] | [0.3715, 0.422, 0.4594] | The word horseradish is attested in English from the 1590s... |
| [white butterfly, pest horseradish, mature cat...] | [0.4587, 0.4974, 0.6484] | Widely introduced by accident... |
我尝试过对words列做索引操作,但只能在序列内部遍历,没达到预期效果。
可复现示例代码:
import pandas as pd from keybert import KeyBERT kw_model = KeyBERT() text = ["The word horseradish is attested in English from the 1590s...","Widely introduced by accident..."] df = pd.DataFrame(text, columns=['texts']) df['words'] = kw_model.extract_keywords(df['texts'], keyphrase_ngram_range=(1, 2), stop_words='english', use_maxsum=True, nr_candidates=20, top_n=3)
解决方案
可以通过pandas的apply方法配合列表推导式,快速拆分元组序列为关键词列表和分数列表:
方法一:直接生成新列
# 先保留原始元组列的副本 df['temp_tuples'] = df['words'].copy() # 提取关键词列表替换原words列 df['words'] = df['temp_tuples'].apply(lambda x: [item[0] for item in x]) # 提取分数列表生成新列 df['scores'] = df['temp_tuples'].apply(lambda x: [item[1] for item in x]) # 删除临时列 df.drop('temp_tuples', axis=1, inplace=True) # 调整列顺序为目标格式 df = df[['texts', 'words', 'scores']]
方法二:用临时DataFrame合并(更简洁)
# 对每个元组序列生成包含words和scores的Series temp_df = df['words'].apply( lambda x: pd.Series({ 'words': [item[0] for item in x], 'scores': [item[1] for item in x] }) ) # 合并原文本列和临时DataFrame df = pd.concat([df['texts'], temp_df], axis=1)
执行以上任意一种方法后,你的DataFrame就会变成目标格式:每行的words列是对应关键词的列表,scores列是匹配的分数列表。
内容的提问来源于stack exchange,提问作者SlowBear
相关产品推荐
相关产品推荐

