如何用Python和Pandas分组统计DataFrame并提取Top3关键词?
实现方法
完全可以用Python + Pandas实现这个需求,核心是通过分组聚合完成统计和文本词频分析,具体步骤如下:
1. 导入依赖库并准备数据
import pandas as pd from collections import Counter import string from nltk.corpus import stopwords import nltk # 下载英文停用词(首次运行需执行) nltk.download('stopwords') stop_words = set(stopwords.words('english')) # 构造示例DataFrame data = { 'ID': [1,2,3,4,5], 'name_x': ['xx','xy','xz','xu','xi'], 'st': ['us','us1','us','us2','us1'], 'string': [ 'Being unacquainted with the chief raccoon was harming his prospects for promotion', 'The overpass went under the highway and into a secret world', 'He was 100% into fasting with her until he understood that meant he couldn\'t eat', 'Random words in front of other random words create a random sentence', 'All you need to do is pick up the pen and begin' ] } df = pd.DataFrame(data)
2. 定义文本处理与词频统计函数
这个函数负责清洗文本、过滤无意义词汇,并返回每组文本中频率最高的前3个关键词(不足3个时补空字符串):
def get_top3_words(text_series): # 合并组内所有文本内容 combined_text = ' '.join(text_series.astype(str)) # 去除标点并转小写,统一词汇格式 translator = str.maketrans('', '', string.punctuation) clean_text = combined_text.lower().translate(translator) # 分词 words = clean_text.split() # 过滤停用词(如the、was)和纯数字 filtered_words = [word for word in words if word not in stop_words and not word.isdigit()] # 统计词频 word_counts = Counter(filtered_words) # 提取前3个高频词,不足3个补空 top_words = [word for word, _ in word_counts.most_common(3)] while len(top_words) < 3: top_words.append('') return top_words
3. 分组聚合生成结果
通过groupby按st分组,同时完成name_x计数和关键词提取:
# 分组处理每个组 result = df.groupby('st').apply( lambda group: pd.Series({ 'name_x_count': group['name_x'].count(), 'top1_word': get_top3_words(group['string'])[0], 'top2_word': get_top3_words(group['string'])[1], 'top3_word': get_top3_words(group['string'])[2] }) ).reset_index() # 查看最终结果 print(result)
结果示例
运行后会得到符合要求的DataFrame:
| st | name_x_count | top1_word | top2_word | top3_word |
|---|---|---|---|---|
| us | 2 | into | was | fasting |
| us1 | 2 | and | the | overpass |
| us2 | 1 | random | words | create |
注:如果想更精准处理缩写(比如couldn't),可以将分词逻辑替换为nltk.word_tokenize,基础split已能满足大部分场景需求。
内容的提问来源于stack exchange,提问作者dd99
相关产品推荐
相关产品推荐

