You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python正则表达式提取文本中2至2,000,000,000的美元金额?

Fixing Regex to Extract Valid USD/US Amounts (2 to 2.24B)

Got it, let's tackle this regex problem for you. First, let's figure out why your original regex ^\{USD}?(\d*(\d\.?|\.\d{1,2}))$ wasn't working:

  • It uses ^ and $ which match the start/end of an entire line, but your amounts are embedded in larger text.
  • \{USD} tries to match literal curly brackets around "USD", which isn't present in your examples.
  • It doesn't account for the "US" prefix, thousand separators (commas), or the required amount range.

Step 1: Build a Regex to Match US/USD + Amount Format

First, let's create a regex that correctly identifies "US" or "USD" followed by a numeric amount (with optional commas and 1-2 decimal places):

\b(US(?:D)?)\s+([\d,]+(?:\.\d{1,2})?)

Let's break this down:

  • \b: Word boundary to avoid matching partial words like "USDA" or "USX".
  • (US(?:D)?): Captures either "US" or "USD" (the (?:D)? is a non-capturing group for the optional "D").
  • \s+: Matches one or more spaces between the prefix and amount (adjust to \s* if there might be no space, e.g., "USD1234").
  • ([\d,]+(?:\.\d{1,2})?): Captures the numeric amount:
    • [\d,]+: Matches digits with optional thousand separators.
    • (?:\.\d{1,2})?: Optional decimal part with 1-2 digits (non-capturing group to keep the clean match).

Step 2: Filter by Amount Range (2 to 2,240,000,000)

Regex isn't great for complex numeric range checks—it gets messy fast. Instead, extract the matches first, then convert them to numbers to validate the range.

Here's a Python example to put it all together:

import re

# Sample input text
input_text = """
Here are some amounts to check:
USD 2,00
US 2,300,000
USD 1.99 (too low)
US 2240000000 (max valid)
USD 2250000000 (too high)
US 500.50
"""

# Our regex pattern
amount_pattern = r'\b(US(?:D)?)\s+([\d,]+(?:\.\d{1,2})?)'

# Extract all potential matches
raw_matches = re.findall(amount_pattern, input_text)

# Process and filter valid amounts
valid_amounts = []
for prefix, amount_str in raw_matches:
    # Remove commas and convert to float
    clean_amount = float(amount_str.replace(',', ''))
    # Check if within the valid range
    if 2 <= clean_amount <= 2240000000:
        # Reconstruct the original formatted string
        valid_amounts.append(f"{prefix} {amount_str}")

print("Valid extracted amounts:")
for item in valid_amounts:
    print(f"- {item}")

Output:

Valid extracted amounts:
- USD 2,00
- US 2,300,000
- US 2240000000
- US 500.50

Notes:

  • This handles odd formats like "USD 2,00" (even though the comma placement looks off—if that's a typo, you could add extra logic to clean up invalid comma positions, but the regex will still capture it).
  • If your text has amounts with no spaces (e.g., "USD123,456"), just change \s+ to \s* in the regex to allow zero or more spaces.

内容的提问来源于stack exchange,提问作者Vanj

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 08:41:06