You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Regex批量分类关键词文本匹配:如何实现出现次数字典统计

文档关键词提取优化需求

场景说明

  • 处理数千份文档(总大小可达2GB)
  • 需匹配约20万个按类别聚合的关键词

当前逐个搜索关键词的方式效率极低,尝试通过管道符编译正则表达式按类别批量匹配,但输出结果不符合预期,需要调整代码以得到目标格式的统计结果。


现有实现代码

import re

text = """
Contrary to popular belief, Lorem Ipsum is not simply random text.
It has roots in a piece of classical Latin literature from 45 BC,
making it over 2000 years old.
Richard McClintock, a Latin professor at Hampden-Sydney College in Virginia,
looked up one of the more obscure Latin words,
consectetur, from a Lorem Ipsum passage,
and going through the cites of the word in classical literature,
discovered the undoubtable source.
Lorem Ipsum comes from sections 1.10.32 and 1.10.33 of
"de Finibus Bonorum et Malorum" (The Extremes of Good and Evil) by Cicero,
written in 45 BC. This book is a treatise on the theory of ethics,
very popular during the Renaissance. The first line of Lorem Ipsum,
"Lorem ipsum dolor sit amet..", comes from a line in section 1.10.32. 
"""

regexes = [
    r'(?P<Writing__book>book)',
    r'(?P<Writing__word>word)',
    r'(?P<Writing__latin>latin)',
    r'(?P<Writing__text>text)',
    r'(?P<Writing__literature>literature)',
    r'(?P<Cities__virginia>virginia)',
    r'(?P<Genre__classical>classical)',
    r'(?P<Genre__renaissance>renaissance)',
]
compiled_regex = '|'.join(regexes)
results = re.findall(
        compiled_regex,
        text,
        flags=re.MULTILINE | re.IGNORECASE
    )
for result in results:
    print(result)

当前运行输出

('', '', '', 'text', '', '', '', '')
('', '', '', '', '', '', 'classical', '')
('', '', 'Latin', '', '', '', '', '')
('', '', '', '', 'literature', '', '', '')
('', '', 'Latin', '', '', '', '', '')
('', '', '', '', '', 'Virginia', '', '')
('', '', 'Latin', '', '', '', '', '')
('', 'word', '', '', '', '', '', '')
('', 'word', '', '', '', '', '', '')
('', '', '', '', '', '', 'classical', '')
('', '', '', '', 'literature', '', '', '')
('book', '', '', '', '', '', '', '')
('', '', '', '', '', '', '', 'Renaissance')

期望输出格式

{'Writing__book': 1, 'Writing__word': 2, 'Cities__virginia': 1, ...}

优化解决方案

问题核心

re.findall在使用命名捕获组的正则时,会返回包含所有捕获组结果的元组,大部分元素为空字符串,无法直接对应到目标键值对。改用re.finditer可以直接获取匹配对象的命名组信息,更高效地统计次数。

优化后代码

import re
from collections import defaultdict

text = """
Contrary to popular belief, Lorem Ipsum is not simply random text.
It has roots in a piece of classical Latin literature from 45 BC,
making it over 2000 years old.
Richard McClintock, a Latin professor at Hampden-Sydney College in Virginia,
looked up one of the more obscure Latin words,
consectetur, from a Lorem Ipsum passage,
and going through the cites of the word in classical literature,
discovered the undoubtable source.
Lorem Ipsum comes from sections 1.10.32 and 1.10.33 of
"de Finibus Bonorum et Malorum" (The Extremes of Good and Evil) by Cicero,
written in 45 BC. This book is a treatise on the theory of ethics,
very popular during the Renaissance. The first line of Lorem Ipsum,
"Lorem ipsum dolor sit amet..", comes from a line in section 1.10.32. 
"""

regexes = [
    r'(?P<Writing__book>book)',
    r'(?P<Writing__word>word)',
    r'(?P<Writing__latin>latin)',
    r'(?P<Writing__text>text)',
    r'(?P<Writing__literature>literature)',
    r'(?P<Cities__virginia>virginia)',
    r'(?P<Genre__classical>classical)',
    r'(?P<Genre__renaissance>renaissance)',
]
compiled_regex = '|'.join(regexes)

# 初始化统计字典,默认值为0
count_dict = defaultdict(int)
# 遍历所有匹配结果
for match in re.finditer(compiled_regex, text, flags=re.MULTILINE | re.IGNORECASE):
    # 获取当前匹配到的命名组键
    matched_key = match.lastgroup
    if matched_key:
        count_dict[matched_key] += 1

# 转换为普通字典输出
result_dict = dict(count_dict)
print(result_dict)

优化后输出

{'Writing__text': 1, 'Genre__classical': 2, 'Writing__latin': 3, 'Writing__literature': 2, 'Cities__virginia': 1, 'Writing__word': 2, 'Writing__book': 1, 'Genre__renaissance': 1}

大规模场景补充建议

  1. 替换正则工具:如果关键词数量达到20万,正则拼接会导致性能瓶颈,推荐使用ahocorasick库(多模式匹配算法),比正则更适合大规模关键词匹配场景。
  2. 文档分块处理:针对2GB级别的文档,建议分块读取,避免一次性加载整个文档占用过多内存。
  3. 大小写策略调整:当前使用re.IGNORECASE统一忽略大小写,若关键词有大小写区分需求,可单独对关键词做预处理或调整正则规则。

内容的提问来源于stack exchange,提问作者Loïc

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 19:30:55