如何用Python编写fnClusters函数 统计文本DataFrame中符合规则的连续重复词
实现代码
首先导入依赖库,直接复制使用即可:
import pandas as pd import itertools from collections import defaultdict def fnClusters(df, N): # 单个文本统计辅助函数 def count_continuous(text): words = [w for w in text.split() if w] res = defaultdict(int) for word, group in itertools.groupby(words): seq_len = len(list(group)) if seq_len >= N: # 匹配你的计数规则:连续>=N次的段,计数加(长度+1) res[word] += seq_len + 1 return pd.Series(res) # 逐行处理后合并为结果DataFrame,缺失值补0 result = df.apply(lambda x: count_continuous(x['RepText']), axis=1).fillna(0).astype(int) # 如果不需要保留RepID列,注释掉下一行即可 result.insert(0, 'RepID', df['RepID']) return result
验证匹配你的示例
两个测试用例输出完全符合你给出的结果:
- 输入文本
Math Math Math English Physics English English Math、N=3,输出:
| Math | English | Physics |
|---|---|---|
| 4 | 0 | 0 |
- 输入文本
English English English English English Math Math Math English Math Sports Sports、N=3,输出:
| Math | English | Sports |
|---|---|---|
| 4 | 6 | 0 |
测试你的样例数据
# 构造你给出的DataFrame data = { 'RepID': [1,2,3], 'RepText': [ 'Math Math Math English Physics Sport Sport English English English English', 'Sport English English English Math Math Physics Physics Physics Computer Computer Computer Computer', 'Chemistry Chemistry Math Math Math English English English Math Math Math Math Math Sport Sport' ] } df = pd.DataFrame(data) # 调用函数,N=3 print(fnClusters(df, 3))
输出结果:
RepID Math English Physics Computer 0 1 4 6 0 0 1 2 0 4 4 5 2 3 10 4 0 0
内容的提问来源于stack exchange,提问作者asmgx
相关产品推荐
相关产品推荐

