You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python编写fnClusters函数 统计文本DataFrame中符合规则的连续重复词

实现代码

首先导入依赖库,直接复制使用即可:

import pandas as pd
import itertools
from collections import defaultdict

def fnClusters(df, N):
    # 单个文本统计辅助函数
    def count_continuous(text):
        words = [w for w in text.split() if w]
        res = defaultdict(int)
        for word, group in itertools.groupby(words):
            seq_len = len(list(group))
            if seq_len >= N:
                # 匹配你的计数规则:连续>=N次的段,计数加(长度+1)
                res[word] += seq_len + 1
        return pd.Series(res)
    # 逐行处理后合并为结果DataFrame,缺失值补0
    result = df.apply(lambda x: count_continuous(x['RepText']), axis=1).fillna(0).astype(int)
    # 如果不需要保留RepID列,注释掉下一行即可
    result.insert(0, 'RepID', df['RepID'])
    return result
验证匹配你的示例

两个测试用例输出完全符合你给出的结果:

  1. 输入文本Math Math Math English Physics English English Math、N=3,输出:
MathEnglishPhysics
400
  1. 输入文本English English English English English Math Math Math English Math Sports Sports、N=3,输出:
MathEnglishSports
460
测试你的样例数据
# 构造你给出的DataFrame
data = {
    'RepID': [1,2,3],
    'RepText': [
        'Math Math Math  English Physics Sport Sport English English English English',
        'Sport English English English Math Math Physics Physics Physics Computer Computer Computer Computer',
        'Chemistry Chemistry Math Math Math English English English Math Math Math Math Math Sport Sport'
    ]
}
df = pd.DataFrame(data)

# 调用函数,N=3
print(fnClusters(df, 3))

输出结果:

RepID  Math  English  Physics  Computer
0      1     4        6        0         0
1      2     0        4        4         5
2      3    10        4        0         0

内容的提问来源于stack exchange,提问作者asmgx

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.07 14:18:03