You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

统计Pandas DataFrame列中指定完整单词的行级出现次数

嘿,这个需求很清晰,我来给你一个简洁高效的解决方案!

解决方案

核心思路

咱们要统计的是完整单词one、two、three的出现次数,得避开像throne这种包含目标字符串但不是独立单词的情况。这里用正则表达式的单词边界\b就能精准匹配独立单词,不会误判。

具体代码实现

import pandas as pd

# 先把你的数据构建成DataFrame(如果已经有现成的DataFrame可以跳过这步)
data = {
    'texts': [
        'throne one',
        'bar one',
        'foo two',
        'bar three',
        'foo two',
        'bar two',
        'foo one',
        'foo three',
        'one three'
    ]
}
df = pd.DataFrame(data)

# 用正则定义要匹配的目标单词,\b确保匹配的是独立单词
target_pattern = r'\b(one|two|three)\b'

# 给DataFrame新增counts列,统计每行的匹配次数
df['counts'] = df['texts'].str.count(target_pattern)

# 查看最终结果
print(df)

结果说明

  • \b(one|two|three)\b这个正则表达式的作用是:只匹配独立存在的one、two或three单词,\b会确保这些单词的前后是单词边界(比如空格、字符串开头/结尾),所以throne里的one不会被错误统计。
  • str.count()方法会自动遍历每行的texts内容,统计符合正则规则的次数,完全贴合你的需求。

运行代码后,你会得到和预期完全一致的输出:

texts  counts
0   throne one       1
1      bar one       1
2      foo two       1
3    bar three       1
4      foo two       1
5      bar two       1
6      foo one       1
7    foo three       1
8    one three       2

内容的提问来源于stack exchange,提问作者StatguyUser

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 08:08:18