You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

大数据集下正则提取失效:小数据集可用的Regex在大数据集不生效

Fixing Your Regex for Larger Datasets

Hey there! Let's break down why your regex works on small test cases but breaks with bigger datasets, and get it sorted out.

The Critical Flaw in Your Current Regex

Your pattern r'Taxi* ([0-9.]+)' has a tiny but impactful mistake: the * is applied to the last character of Taxi (the i), not the whole word. That means it matches strings like Tax (0 is), Taxi (1 i), Taxii (2 is), etc.—but it only looks for a direct space followed by numbers right after that match.

In your small dataset, you probably had consistent formatting where Taxi was immediately followed by a space and the number (even though your example shows Taxi ID: 170618970—maybe that was a one-off or the regex accidentally lined up?). But larger datasets almost always have varied formatting: think Taxi number: 12345, taxi ref 678.90, TAXI ID: 000123, etc. Your original regex can't handle any of these cases because it's way too rigid.

The Fixes

We need to adjust the regex to be flexible enough to handle whatever text sits between Taxi and your target number, while still reliably capturing the digits. Here are two solid options:

Option 1: Flexible Match for Non-Digit Text

Use this regex to match Taxi (case-insensitive), then any non-digit characters (spaces, colons, labels like "ID:"), then capture the number:

r'Taxi\D*([0-9.]+)'
  • Taxi: Matches the exact word (we removed the incorrect * here)
  • \D*: Matches 0 or more non-digit characters (covers spaces, labels, punctuation—anything between Taxi and the number)
  • ([0-9.]+): Captures your target number (including decimals if needed)

When you run this with re.IGNORECASE on your example string Goa return Taxi ID: 170618970, it will correctly capture 170618970.

Option 2: Target Specific Labels (More Precise)

If you know the common labels that appear between Taxi and the number (like "ID:", "number:", "ref:"), you can make the regex more precise to avoid accidental matches:

r'Taxi(?:\s+(?:ID|number|ref):)?\s*([0-9.]+)'
  • (?:\s+(?:ID|number|ref):)?: Optional non-capturing group for common labels (so it matches Taxi ID:, Taxi number:, or just Taxi)
  • \s*: Matches any extra spaces before the number

Testing It Out

Try running either of these regexes against your larger dataset—they should handle all the varied formatting that broke your original pattern. Remember to keep using re.IGNORECASE to account for uppercase/lowercase variations like TAXI or taxi.

内容的提问来源于stack exchange,提问作者Anubhay Jha

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 04:17:57