You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中按平台统计DataFrame内Hashtag出现次数时遭遇AttributeError的技术求助

解决按平台统计Hashtag出现次数的问题

嘿,我来帮你搞定这个报错!你遇到的AttributeError: 'SeriesGroupBy' object has no attribute 'str'是因为当你用groupby(['Post', 'Platform'])['Post']后得到的是SeriesGroupBy对象,这个类型并没有内置的str属性,所以没法直接调用str.extractall来提取hashtag。

下面给你两种简洁可行的解决方案,都能实现按平台统计每个hashtag的出现次数:

方法一:先提取所有Hashtag再关联平台信息

这种方法先把所有帖子里的hashtag提取出来,再通过索引关联对应的平台,最后分组统计:

# 提取所有hashtag,同时保留原数据的索引(level_0)
hashtags_extracted = df['Post'].str.extractall(r'(\#\w+)').reset_index()
# 关联原数据的Platform列,并重命名hashtag列
hashtags_with_platform = hashtags_extracted.merge(
    df[['Platform']], 
    left_on='level_0', 
    right_index=True
).rename(columns={0: 'hashtags'})
# 按平台和hashtag分组统计次数
result = hashtags_with_platform.groupby(['Platform', 'hashtags']).size().reset_index(name='count')

方法二:用findall+explode拆分Hashtag

这种方法先把每个帖子的hashtag转换成列表,再拆分成单独的行,之后直接分组统计:

# 把每个帖子里的hashtag提取为列表
df['hashtags'] = df['Post'].str.findall(r'(\#\w+)')
# 将列表拆分成多行,每个hashtag占一行
df_exploded = df.explode('hashtags')
# 按平台和hashtag分组统计次数
result = df_exploded.groupby(['Platform', 'hashtags']).size().reset_index(name='count')

用你的示例数据运行后,两种方法都会得到如下结果:

Platformhashtagscount
Insta#hashtag21
Twitter#hashtag12

内容的提问来源于stack exchange,提问作者DotPi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 08:39:11