You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas按年份分组统计文本列中各关键词对应行数的实现方法

Pandas按年统计关键词出现行数解决方案

实现逻辑如下:先给每行生成是否包含各关键词的0/1标识列,再按年份维度分组求和即可,代码如下:

import pandas as pd

# 你的示例数据
txt = """2019  'an example sentence'
2019  'another sentence'
2020  'fox trot'
2020  'this sentence has some new words'
2020  'new and old example words here'
2021  'the fox jumps for example'
2021  'a different example sentence'"""

df = pd.DataFrame([x.split('  ') for x in txt.split('\n')], columns=['year','text'])
keywords = ['example', 'sentence', 'fox']

# 核心实现代码
# 1. 年份转整数类型,避免后续分组、排序异常
df['year'] = df['year'].astype(int)

# 2. 生成各关键词的匹配标识列(1代表当前行包含该关键词,0代表不包含)
for kw in keywords:
    # 如果关键词包含正则特殊字符,需加regex=False参数避免匹配错误
    df[kw] = df['text'].str.contains(kw).astype(int)

# 3. 按年份分组,对各关键词列求和得到年度统计结果
result = df.groupby('year', as_index=False)[keywords].sum()

print(result)

注意:如果你的关键词包含.、*、?这类正则特殊字符,需要将str.contains(kw)修改为str.contains(kw, regex=False),避免出现非预期的匹配结果。

内容的提问来源于stack exchange,提问作者End genocide - save Gaza

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 22:09:02