如何从DataFrame中提取列表单词出现频次最高的Top N行?
问题描述
我有一个包含多行多列的Python DataFrame,示例数据如下:
| DocName | Content |
|---|---|
| Doc1 | Hi how you are doing ? Hope you are well. I hear the food is great! |
| Doc2 | The food is great. James loves his food. You not so much right ? |
| Doc3. | Yeah he is alright. |
我还有一个包含100个单词的列表,格式如下:
word_list = ["food", "you", ....]
现在需要从Content列中,提取列表中每个单词出现频次最高的前N行。针对上述示例数据:
"food"在Doc2中出现2次,在Doc1中出现1次。
"you"在Doc1中出现2次,在Doc2中出现1次。
期望输出为:
{"food": ["Doc2", "Doc1"], "you": ["Doc1", "Doc2"], .....}
其中N=2(即每个单词对应出现频次最高的前2行)
我尝试了以下代码,但不确定后续如何推进:
word_list = ["food", "you", ....] result = [] for word in word_list: result.append(df.Content.apply(lambda row: sum([row.count(word)])))
请问如何在Python中实现满足上述需求的高效解决方案?
高效解决方案
基础实现(适合中小规模数据)
先统一文本大小写避免计数偏差,再用向量化操作计算词频并排序:
import pandas as pd from collections import defaultdict # 示例DataFrame data = { "DocName": ["Doc1", "Doc2", "Doc3."], "Content": [ "Hi how you are doing ? Hope you are well. I hear the food is great!", "The food is great. James loves his food. You not so much right ?", "Yeah he is alright." ] } df = pd.DataFrame(data) # 统一文本为小写 df['Content_lower'] = df['Content'].str.lower() def get_top_docs(df, word_list, top_n=2): result = defaultdict(list) for word in word_list: # 向量化计算每个文档中当前单词的出现次数 word_counts = df['Content_lower'].str.count(word.lower()) # 按频次降序获取文档名称的排序结果 ranked_docs = df['DocName'].iloc[word_counts.argsort(ascending=False)] # 取前N个文档存入结果 result[word] = ranked_docs[:top_n].tolist() return dict(result) # 调用函数 word_list = ["food", "you"] top_n = 2 final_result = get_top_docs(df, word_list, top_n) print(final_result) # 输出:{'food': ['Doc2', 'Doc1'], 'you': ['Doc1', 'Doc2']}
代码说明
- 使用
str.count()向量化操作替代逐行循环,效率更高 argsort(ascending=False)获取频次降序的索引,再通过iloc匹配对应的文档名称- 用
defaultdict简化字典的初始化和赋值逻辑
优化实现(适合大规模单词/数据)
如果单词数量多(比如100个)或DataFrame规模大,推荐用CountVectorizer一次性计算所有目标单词的词频矩阵,避免多次遍历数据:
import pandas as pd from sklearn.feature_extraction.text import CountVectorizer # 示例DataFrame data = { "DocName": ["Doc1", "Doc2", "Doc3."], "Content": [ "Hi how you are doing ? Hope you are well. I hear the food is great!", "The food is great. James loves his food. You not so much right ?", "Yeah he is alright." ] } df = pd.DataFrame(data) word_list = ["food", "you"] top_n = 2 # 初始化CountVectorizer,仅保留目标单词 vec = CountVectorizer(vocabulary=word_list, lowercase=True) # 生成词频矩阵(行=文档,列=单词) count_matrix = vec.fit_transform(df['Content']) # 转换为DataFrame,索引为文档名称 count_df = pd.DataFrame( count_matrix.toarray(), index=df['DocName'], columns=vec.get_feature_names_out() ) # 遍历每个单词,提取频次最高的前N个文档 result = {} for word in word_list: top_docs = count_df[word].sort_values(ascending=False).index[:top_n].tolist() result[word] = top_docs print(result) # 输出:{'food': ['Doc2', 'Doc1'], 'you': ['Doc1', 'Doc2']}
优化说明
- 一次性完成所有目标单词的词频计算,减少数据遍历次数
- 利用
sklearn的高效文本处理能力,适合大规模场景
内容的提问来源于stack exchange,提问作者dravid07
相关产品推荐
相关产品推荐

