如何在正则中忽略特定匹配?PDF审计师提取正则优化求助
解决审计师名称提取的正则干扰问题
我来帮你搞定这个问题——核心是要排除包含issued by the Hong Kong Institute of的干扰匹配,同时不用那种可能误杀合法结果的固定前缀否定。这里有两种可靠的方案:
方案1:用否定环视优化正则(Python 3.6+兼容)
我们可以给正则加一个否定前瞻断言,确保匹配到的审计师行后面不会跟着那段干扰文本:
pattern = r'.*\n.*?(?P<auditor>[A-Z].+?)(?:LLP\s*)?\s*((PRC.*?|Chinese.*?)?[Cc]ertified [Pp]ublic|[Cc]hartered) [Aa]ccountants(?!.*issued by the Hong Kong Institute of)'
这个正则的逻辑是:匹配符合审计师格式的内容,但同时确保该内容之后的文本里不包含干扰字符串,从根源上排除无效匹配。
方案2:先收集候选再过滤(兼容所有Python版本,更稳妥)
如果要兼容更早的Python版本,或者想让逻辑更清晰、更易维护,我们可以先提取所有符合格式的候选结果,再手动过滤掉包含干扰文本的项:
修改后的完整代码
from io import BytesIO import pdfplumber, requests, re test_case = { 'https://www1.hkexnews.hk/listedco/listconews/sehk/2020/0514/2020051400555.pdf': 59, 'https://www1.hkexnews.hk/listedco/listconews/gem/2020/0529/2020052902118.pdf': 55, 'https://www1.hkexnews.hk/listedco/listconews/sehk/2020/0618/2020061800366.pdf': 47, 'https://www1.hkexnews.hk/listedco/listconews/gem/2020/0630/2020063002674.pdf': 30, } # 先匹配所有符合审计师格式的候选,不做严格前缀限制 pattern = r'.*\n.*?(?P<auditor>[A-Z].+?)(?:LLP\s*)?\s*((PRC.*?|Chinese.*?)?[Cc]ertified [Pp]ublic|[Cc]hartered) [Aa]ccountants' for url, page in test_case.items(): rq = requests.get(url) pdf = pdfplumber.load(BytesIO(rq.content)) txt = pdf.pages[page].extract_text() txt = re.sub("([^\x00-\x7F])+", "", txt) # 移除中文 # 找到所有匹配的候选项 matches = re.finditer(pattern, txt, flags=re.MULTILINE) valid_auditor = None for match in matches: candidate = match.group('auditor').strip() # 检查当前匹配的完整文本是否包含干扰字符串 full_match_text = match.group(0) if 'issued by the Hong Kong Institute of' not in full_match_text: valid_auditor = candidate break # 找到第一个有效匹配就停止 if valid_auditor: print(repr(valid_auditor)) else: print(txt) print('============') print(url)
为什么这个方案更优?
- 通用性强:不会像
^(?!Hong|Kong)那样,误过滤掉名字里包含这些词的合法审计师(比如如果有个审计师叫Kong & Partners,之前的正则就会漏匹配)。 - 逻辑清晰:先收集所有可能的结果,再过滤干扰项,后续如果有新的干扰文本,只需要修改过滤条件即可,维护成本更低。
运行结果
两种方案都能输出你期望的结果:
'ShineWing' 'ShineWing' 'Ernst & Young' 'Elite Partners CPA Limited'
内容的提问来源于stack exchange,提问作者chris_2020
相关产品推荐
相关产品推荐

