You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Pandas DataFrame中按subreddit分组统计分词词频

按Subreddit分组获取高频分词的解决方案

没问题,我来帮你搞定这个需求!要实现按subreddit列分组,提取每组tokenized_text里的高频分词,咱们可以结合Pandas和Python内置的Counter来完成,步骤如下:

1. 导入所需工具库

首先得把需要的库导入进来,pandas用来处理DataFrame,Counter做词频统计,chain用来扁平化分词列表:

import pandas as pd
from collections import Counter
from itertools import chain

2. 定义词频统计函数

咱们写一个自定义函数,用来处理每个分组的分词数据:

def get_top_tokens(token_series, top_n=10):
    # 把分组里的所有分词列表合并成一个大列表
    all_tokens = list(chain.from_iterable(token_series))
    # 统计每个分词出现的频率
    token_counts = Counter(all_tokens)
    # 返回出现次数最多的前N个分词(默认取前10)
    return token_counts.most_common(top_n)

这个函数的作用是:接收一个分组后的分词序列,先把所有分散的分词列表合并成一个整体,然后统计每个词的出现次数,最后返回排序好的高频词。

3. 应用到分组数据上

接下来把这个函数应用到按subreddit分组后的DataFrame上:

# 按subreddit分组,对tokenized_text列应用统计函数
top_tokens_per_subreddit = df.groupby('subreddit')['tokenized_text'].apply(get_top_tokens)

4. 查看结果

现在你就可以查看每个subreddit的高频分词了,比如:

# 查看前5个subreddit的高频词
print(top_tokens_per_subreddit.head())

# 单独查看'15SecondStories'板块的Top 10高频词
print(top_tokens_per_subreddit['15SecondStories'])

可选优化:过滤停用词

如果想去掉像"the"、"is"这类无意义的停用词,可以借助NLTK的停用词库来优化函数:
首先得下载停用词并导入:

import nltk
nltk.download('stopwords')
from nltk.corpus import stopwords

stop_words = set(stopwords.words('english'))

然后修改统计函数:

def get_top_tokens(token_series, top_n=10):
    all_tokens = list(chain.from_iterable(token_series))
    # 过滤停用词和长度小于2的词(可选)
    filtered_tokens = [token for token in all_tokens if token not in stop_words and len(token) > 1]
    token_counts = Counter(filtered_tokens)
    return token_counts.most_common(top_n)

这样得到的高频词会更有实际意义~

内容的提问来源于stack exchange,提问作者Parseltongue

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 11:25:01