正则表达式语法错误修复:提取<PAGE>后指定内容的问题解决
正则表达式报错排查与修正方案
错误原因分析
我帮你拆解下触发TypeError: 'str' object is not callable的核心问题:
- 语法结构错误:你把正则表达式拆成了多个字符串加括号的形式,比如
('?<=<PAGE>')('.*')('?=Management\'s Discussion'),Python会把第一个字符串'(?<=<PAGE>)'当成可调用对象,后面的('.*')就变成了调用这个字符串,而字符串本身不能被调用,所以直接抛出类型错误。 - 正则断言语法错误:正确的后行断言语法是
(?<=...),先行断言是(?=...),你漏写了括号里的<和>,而且拆分的写法完全不符合Python正则的定义规则。 - 大小写不匹配:原文中的目标文本是全大写的
MANAGEMENT'S DISCUSSION,但你写的是Management\'s Discussion,大小写不一致会导致匹配失败。
修正后的代码方案
如果你需要提取<PAGE>之后出现的MANAGEMENT'S DISCUSSION AND ANALYSIS,或者匹配<PAGE>到该文本之间的内容,这里给出两种场景的修正代码:
场景1:直接提取<PAGE>之后的MANAGEMENT'S DISCUSSION AND ANALYSIS
import re # 你的目标文本 target_text = """(3) Reflects the adoption of SFAS No. 128, EARNINGS PER SHARE. <PAGE> ITEM 7. MANAGEMENT'S DISCUSSION AND ANALYSIS OF FINANCIAL CONDITION AND RESULTS OF OPERATION YEAR ENDED DECEMBER 28, 1997 COMPARED TO THE YEAR ENDED DECEMBER 29, 1996 In November 1996, the Company initiated a major restructuring and growth plan designed to substantially reduce its cost structure and grow the business in order to restore higher levels of profitability for the Company. By July 1997, the Company completed the major phases of the restructuring plan. The $225.0 million of annualized cost savings anticipated from the restructuring results primarily from the consolidation of administrative functions within """ # 修正后的正则:用原始字符串避免转义问题,添加忽略大小写标志兼容全大写文本 pattern = r'(?<=<PAGE>.*)MANAGEMENT\'S DISCUSSION AND ANALYSIS' matches = re.findall(pattern, target_text, re.IGNORECASE) print(matches) # 输出: ["MANAGEMENT'S DISCUSSION AND ANALYSIS"]
场景2:匹配<PAGE>到MANAGEMENT'S DISCUSSION之间的内容
import re target_text = """(3) Reflects the adoption of SFAS No. 128, EARNINGS PER SHARE. <PAGE> ITEM 7. MANAGEMENT'S DISCUSSION AND ANALYSIS OF FINANCIAL CONDITION AND RESULTS OF OPERATION YEAR ENDED DECEMBER 28, 1997 COMPARED TO THE YEAR ENDED DECEMBER 29, 1996 In November 1996, the Company initiated a major restructuring and growth plan designed to substantially reduce its cost structure and grow the business in order to restore higher levels of profitability for the Company. By July 1997, the Company completed the major phases of the restructuring plan. The $225.0 million of annualized cost savings anticipated from the restructuring results primarily from the consolidation of administrative functions within """ # 用非贪婪匹配.*?避免匹配过多内容,确保只取到目标文本之前的部分 pattern = r'(?<=<PAGE>).*?(?=MANAGEMENT\'S DISCUSSION)' matches = re.findall(pattern, target_text, re.IGNORECASE) print(matches) # 输出: [" ITEM 7. "]
关键修正点总结
- 将正则表达式定义为完整的原始字符串(用
r前缀),避免转义字符干扰,同时不要拆分字符串。 - 修复断言语法:正确使用
(?<=...)后行断言和(?=...)先行断言。 - 添加
re.IGNORECASE标志,兼容原文的全大写格式,避免大小写不匹配导致的匹配失败。 - 用
.*?非贪婪匹配替代.*贪婪匹配,确保在存在多个目标文本时不会过度匹配。
内容的提问来源于stack exchange,提问作者user9824386
相关产品推荐
相关产品推荐

