You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从DataFrame中提取列表单词出现频次最高的Top N行?

问题描述

我有一个包含多行多列的Python DataFrame,示例数据如下:

DocNameContent
Doc1Hi how you are doing ? Hope you are well. I hear the food is great!
Doc2The food is great. James loves his food. You not so much right ?
Doc3.Yeah he is alright.

我还有一个包含100个单词的列表,格式如下:

word_list = ["food", "you", ....]           

现在需要从Content列中,提取列表中每个单词出现频次最高的前N行。针对上述示例数据:

"food"在Doc2中出现2次,在Doc1中出现1次。
"you"在Doc1中出现2次,在Doc2中出现1次。

期望输出为:

{"food": ["Doc2", "Doc1"], "you": ["Doc1", "Doc2"], .....}

其中N=2(即每个单词对应出现频次最高的前2行)

我尝试了以下代码,但不确定后续如何推进:

word_list = ["food", "you", ....]
result = []

for word in word_list:
    result.append(df.Content.apply(lambda row: sum([row.count(word)])))

请问如何在Python中实现满足上述需求的高效解决方案?


高效解决方案

基础实现(适合中小规模数据)

先统一文本大小写避免计数偏差,再用向量化操作计算词频并排序:

import pandas as pd
from collections import defaultdict

# 示例DataFrame
data = {
    "DocName": ["Doc1", "Doc2", "Doc3."],
    "Content": [
        "Hi how you are doing ? Hope you are well. I hear the food is great!",
        "The food is great. James loves his food. You not so much right ?",
        "Yeah he is alright."
    ]
}
df = pd.DataFrame(data)

# 统一文本为小写
df['Content_lower'] = df['Content'].str.lower()

def get_top_docs(df, word_list, top_n=2):
    result = defaultdict(list)
    for word in word_list:
        # 向量化计算每个文档中当前单词的出现次数
        word_counts = df['Content_lower'].str.count(word.lower())
        # 按频次降序获取文档名称的排序结果
        ranked_docs = df['DocName'].iloc[word_counts.argsort(ascending=False)]
        # 取前N个文档存入结果
        result[word] = ranked_docs[:top_n].tolist()
    return dict(result)

# 调用函数
word_list = ["food", "you"]
top_n = 2
final_result = get_top_docs(df, word_list, top_n)
print(final_result)
# 输出:{'food': ['Doc2', 'Doc1'], 'you': ['Doc1', 'Doc2']}

代码说明

  • 使用str.count()向量化操作替代逐行循环,效率更高
  • argsort(ascending=False)获取频次降序的索引,再通过iloc匹配对应的文档名称
  • 用defaultdict简化字典的初始化和赋值逻辑

优化实现(适合大规模单词/数据)

如果单词数量多(比如100个)或DataFrame规模大,推荐用CountVectorizer一次性计算所有目标单词的词频矩阵,避免多次遍历数据:

import pandas as pd
from sklearn.feature_extraction.text import CountVectorizer

# 示例DataFrame
data = {
    "DocName": ["Doc1", "Doc2", "Doc3."],
    "Content": [
        "Hi how you are doing ? Hope you are well. I hear the food is great!",
        "The food is great. James loves his food. You not so much right ?",
        "Yeah he is alright."
    ]
}
df = pd.DataFrame(data)

word_list = ["food", "you"]
top_n = 2

# 初始化CountVectorizer,仅保留目标单词
vec = CountVectorizer(vocabulary=word_list, lowercase=True)
# 生成词频矩阵(行=文档,列=单词)
count_matrix = vec.fit_transform(df['Content'])
# 转换为DataFrame,索引为文档名称
count_df = pd.DataFrame(
    count_matrix.toarray(),
    index=df['DocName'],
    columns=vec.get_feature_names_out()
)

# 遍历每个单词,提取频次最高的前N个文档
result = {}
for word in word_list:
    top_docs = count_df[word].sort_values(ascending=False).index[:top_n].tolist()
    result[word] = top_docs

print(result)
# 输出:{'food': ['Doc2', 'Doc1'], 'you': ['Doc1', 'Doc2']}

优化说明

  • 一次性完成所有目标单词的词频计算,减少数据遍历次数
  • 利用sklearn的高效文本处理能力,适合大规模场景

内容的提问来源于stack exchange,提问作者dravid07

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 18:45:39