如何用Python的re模块提取指定短语后的首个数值?
正则表达式提取指定短语后的首个数值问题
问题场景
需要从给定句子中,提取短语**“tangible net worth”之后出现的首个数值**,具体示例如下:
示例句子
- "A company must maintain a minimum tangible net worth of $100000000 and leverage ratio of 0.5"
- "Minimum required tangible net worth the firm needs to maintain is $50000000"
预期目标
从两句中分别提取$100000000和$50000000,最终生成如下字典:
{ "tangible net worth": "$100000000" }
尝试过的无效正则
以下正则表达式均未得到预期结果:
re.search(r'net worth.*(\d+)', sent) re.search(r'(net worth)(.*)(\d+)', sent) re.search(r'(net worth)(.*)(\d?)', sent) re.findall(r'tangible net worth (.*)?(\d* )', sent) re.findall(r'tangible net worth (.*)?( \d* )', sent) re.findall(r'tangible net worth (.*)?(\d)', sent)
解决方法
核心是用非贪婪匹配锁定短语到首个数值的区间,精准捕获目标数值。
有效正则表达式
使用re.search搭配以下正则:
r'tangible net worth.*?(\$\d+)'
tangible net worth:精准匹配目标短语.*?:非贪婪匹配任意字符,确保只取到首个数值前的内容,不会跳过目标去匹配后面的数值(\$\d+):捕获以$开头、后跟连续数字的数值(若需要兼容小数,可改为(\$\d+\.?\d*))
完整Python代码
import re # 待处理的句子列表 sentences = [ "A company must maintain a minimum tangible net worth of $100000000 and leverage ratio of 0.5", "Minimum required tangible net worth the firm needs to maintain is $50000000" ] target_phrase = "tangible net worth" result = {} for sent in sentences: match = re.search(r'tangible net worth.*?(\$\d+)', sent) if match: # 若要收集所有句子的结果,可改为列表存储 result[target_phrase] = match.group(1) print(result)
补充说明
- 如果句子中短语的大小写不固定,可添加
re.IGNORECASE忽略大小写:re.search(r'tangible net worth.*?(\$\d+)', sent, re.IGNORECASE) - 若数值格式多样(比如不带$、带逗号分隔如
$100,000),可调整捕获组为([$]?\d[,\d]*)来适配
内容的提问来源于stack exchange,提问作者Ruchit
相关产品推荐
相关产品推荐

