Python提取两个指定子串间的文本 解决正则匹配返回None报错问题
正则提取问题修复方案
报错原因
AttributeError: 'NoneType' object has no attribute 'group'触发原因是正则没有匹配到任何内容,re.search返回空值,直接调用group方法就会报错。
原代码核心问题
- 匹配前缀写错:目标前缀是
Listen | >,原正则写的是Listen;,和目标文本完全不匹配 - 正则规则错误:
|是正则的特殊字符需要转义,同时默认.*是贪婪匹配,会优先匹配到最远的1字符,不符合提取最近区间的需求
修复后代码
import re text = "1 (2 points) fa] 4) Listen | > Apache Cassandra is an open source NoSQL distributed database that delivers scalability and high availability without compromising performance and is trusted by thousands of companies. Linear scalability and proven fault tolerance on 1) commodity hardward © 2) ubuntu Os) ovals" # r标记表示原生字符串,\|转义竖线,.*?开启非贪婪匹配,\s*自动过滤首尾多余空格 result = re.search(r'Listen \| >\s*(.*?)\s*1', text) if result is not None: # 提取英文结果 extract_content = result.group(1) print(extract_content) # 提取内容中文翻译 # print("提取内容翻译:Apache Cassandra是一款开源NoSQL分布式数据库,在不损失性能的前提下具备可扩展性与高可用性,获得数千家企业的信赖,拥有线性扩展能力,在通用硬件上已被验证具备容错能力")
运行输出
Apache Cassandra is an open source NoSQL distributed database that delivers scalability and high availability without compromising performance and is trusted by thousands of companies. Linear scalability and proven fault tolerance on
内容的提问来源于stack exchange,提问作者John Anthony
相关产品推荐
相关产品推荐

