You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python和Pandas分组统计DataFrame并提取Top3关键词?

实现方法

完全可以用Python + Pandas实现这个需求,核心是通过分组聚合完成统计和文本词频分析,具体步骤如下:

1. 导入依赖库并准备数据

import pandas as pd
from collections import Counter
import string
from nltk.corpus import stopwords
import nltk

# 下载英文停用词(首次运行需执行)
nltk.download('stopwords')
stop_words = set(stopwords.words('english'))

# 构造示例DataFrame
data = {
    'ID': [1,2,3,4,5],
    'name_x': ['xx','xy','xz','xu','xi'],
    'st': ['us','us1','us','us2','us1'],
    'string': [
        'Being unacquainted with the chief raccoon was harming his prospects for promotion',
        'The overpass went under the highway and into a secret world',
        'He was 100% into fasting with her until he understood that meant he couldn\'t eat',
        'Random words in front of other random words create a random sentence',
        'All you need to do is pick up the pen and begin'
    ]
}
df = pd.DataFrame(data)

2. 定义文本处理与词频统计函数

这个函数负责清洗文本、过滤无意义词汇,并返回每组文本中频率最高的前3个关键词(不足3个时补空字符串):

def get_top3_words(text_series):
    # 合并组内所有文本内容
    combined_text = ' '.join(text_series.astype(str))
    # 去除标点并转小写,统一词汇格式
    translator = str.maketrans('', '', string.punctuation)
    clean_text = combined_text.lower().translate(translator)
    # 分词
    words = clean_text.split()
    # 过滤停用词(如the、was)和纯数字
    filtered_words = [word for word in words if word not in stop_words and not word.isdigit()]
    # 统计词频
    word_counts = Counter(filtered_words)
    # 提取前3个高频词,不足3个补空
    top_words = [word for word, _ in word_counts.most_common(3)]
    while len(top_words) < 3:
        top_words.append('')
    return top_words

3. 分组聚合生成结果

通过groupby按st分组,同时完成name_x计数和关键词提取:

# 分组处理每个组
result = df.groupby('st').apply(
    lambda group: pd.Series({
        'name_x_count': group['name_x'].count(),
        'top1_word': get_top3_words(group['string'])[0],
        'top2_word': get_top3_words(group['string'])[1],
        'top3_word': get_top3_words(group['string'])[2]
    })
).reset_index()

# 查看最终结果
print(result)

结果示例

运行后会得到符合要求的DataFrame:

stname_x_counttop1_wordtop2_wordtop3_word
us2intowasfasting
us12andtheoverpass
us21randomwordscreate

注:如果想更精准处理缩写(比如couldn't),可以将分词逻辑替换为nltk.word_tokenize,基础split已能满足大部分场景需求。

内容的提问来源于stack exchange,提问作者dd99

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 03:35:19