You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python正则表达式无法匹配指定文本中目标数值问题求助

问题:提取指定条件下的数值

需要提取「Tax Exempt Securities」区块下「Sales Cost Removed」对应的数值,且仅当该区块出现在「Fixed Income」之前时才提取。目前现有正则表达式可正常匹配「Purchases」对应的数值,但匹配「Sales Cost Removed」或「Other」时失败。

测试文本

Equities
  Purchases
503,773.09
3,900,439.7
  Sales Cost Removed
(397,196.15)
(3,835,270.54)
  Other
-
(2,452.33)
Tax Exempt Securities
  Purchases
301,596.54
606,468.27
  Sales Cost Removed
(350,825.56)
(688,845.64)
  Other
1,268.59
2,000.26
Fixed Income
  Other
-
-

尝试的正则表达式

pattern_1 = re.compile(r"Tax Exempt Securities[\s\S]((?!Fixed Income).)*?Sales Cost Removed[\s\S]*?(\d[\d,]*\.?\d*)")

相关Python代码

for page_num in range(len(pdf_reader.pages)):
    page = pdf_reader.pages[page_num]
    text = page.extract_text()
    print(text)
    print(f"Searching in page {page_num+1}...")
    matches = pattern_1.findall(text)
    if matches:
        print(f"Matches found: {matches}")
        value_to_write = matches[0][1].replace(',', '')
        print(value_to_write)
        try:
            value_to_write = float(value_to_write)
            print(f"Value to write: {value_to_write}")
        except ValueError:
            print("Conversion to float failed, setting value to 0")
            value_to_write = 0
        break
    else:
        print("No match found on this page.")

解决方案

原正则失败原因

  1. 目标数值以括号包裹(如(350,825.56)),但原正则的捕获组(\d[\d,]*\.?\d*)仅匹配以数字开头的内容,无法识别括号开头的数值。
  2. 正则结构冗余,[\s\S]((?!Fixed Income).)*?会多匹配一个字符,导致后续匹配逻辑出错。

修正后的正则表达式

pattern = re.compile(r"Tax Exempt Securities(?:(?!Fixed Income)[\s\S])*?Sales Cost Removed[\s\S]*?\(([\d,]+\.\d+)\)")

正则说明

  • Tax Exempt Securities:准确定位目标数据所在区块的开头
  • (?:(?!Fixed Income)[\s\S])*?:非捕获组,匹配任意字符(含换行),且确保不会越过Fixed Income,保证目标区块在Fixed Income之前
  • Sales Cost Removed:定位到目标数值对应的标签行
  • [\s\S]*?:匹配标签行到数值之间的空白和换行(非贪婪模式)
  • \(([\d,]+\.\d+)\):捕获括号内的数值内容,适配带括号的负数格式

修改后的Python代码示例

import re

# 示例文本(实际使用时替换为PDF提取的text)
text = """Equities
  Purchases
503,773.09
3,900,439.7
  Sales Cost Removed
(397,196.15)
(3,835,270.54)
  Other
-
(2,452.33)
Tax Exempt Securities
  Purchases
301,596.54
606,468.27
  Sales Cost Removed
(350,825.56)
(688,845.64)
  Other
1,268.59
2,000.26
Fixed Income
  Other
-
-"""

pattern = re.compile(r"Tax Exempt Securities(?:(?!Fixed Income)[\s\S])*?Sales Cost Removed[\s\S]*?\(([\d,]+\.\d+)\)")
matches = pattern.findall(text)

if matches:
    print(f"匹配到的数值:{matches}")
    # 处理第一个数值(可根据需求调整为处理全部)
    value_to_write = matches[0].replace(',', '')
    try:
        value_to_write = float(value_to_write)
        # 原数值是括号包裹的负数,转为负浮点数
        value_to_write = -value_to_write
        print(f"处理后的数值:{value_to_write}")
    except ValueError:
        print("转换失败,设置为0")
        value_to_write = 0
else:
    print("未找到匹配内容")

补充说明

如果需要匹配所有Sales Cost Removed下的数值(包括多个),上述正则的findall会返回所有符合条件的结果。若数值存在无括号的正数格式,可进一步调整正则以兼容两种情况:

pattern = re.compile(r"Tax Exempt Securities(?:(?!Fixed Income)[\s\S])*?Sales Cost Removed[\s\S]*?(?:\(([\d,]+\.\d+)\)|([\d,]+\.\d+))")

该正则可同时匹配带括号的负数和无括号的正数,后续需判断两个捕获组哪个有值并做对应处理。

内容的提问来源于stack exchange,提问作者Chance Hunter

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 13:53:21