大数据集下正则提取失效:小数据集可用的Regex在大数据集不生效
Hey there! Let's break down why your regex works on small test cases but breaks with bigger datasets, and get it sorted out.
The Critical Flaw in Your Current Regex
Your pattern r'Taxi* ([0-9.]+)' has a tiny but impactful mistake: the * is applied to the last character of Taxi (the i), not the whole word. That means it matches strings like Tax (0 is), Taxi (1 i), Taxii (2 is), etc.—but it only looks for a direct space followed by numbers right after that match.
In your small dataset, you probably had consistent formatting where Taxi was immediately followed by a space and the number (even though your example shows Taxi ID: 170618970—maybe that was a one-off or the regex accidentally lined up?). But larger datasets almost always have varied formatting: think Taxi number: 12345, taxi ref 678.90, TAXI ID: 000123, etc. Your original regex can't handle any of these cases because it's way too rigid.
The Fixes
We need to adjust the regex to be flexible enough to handle whatever text sits between Taxi and your target number, while still reliably capturing the digits. Here are two solid options:
Option 1: Flexible Match for Non-Digit Text
Use this regex to match Taxi (case-insensitive), then any non-digit characters (spaces, colons, labels like "ID:"), then capture the number:
r'Taxi\D*([0-9.]+)'
Taxi: Matches the exact word (we removed the incorrect*here)\D*: Matches 0 or more non-digit characters (covers spaces, labels, punctuation—anything betweenTaxiand the number)([0-9.]+): Captures your target number (including decimals if needed)
When you run this with re.IGNORECASE on your example string Goa return Taxi ID: 170618970, it will correctly capture 170618970.
Option 2: Target Specific Labels (More Precise)
If you know the common labels that appear between Taxi and the number (like "ID:", "number:", "ref:"), you can make the regex more precise to avoid accidental matches:
r'Taxi(?:\s+(?:ID|number|ref):)?\s*([0-9.]+)'
(?:\s+(?:ID|number|ref):)?: Optional non-capturing group for common labels (so it matchesTaxi ID:,Taxi number:, or justTaxi)\s*: Matches any extra spaces before the number
Testing It Out
Try running either of these regexes against your larger dataset—they should handle all the varied formatting that broke your original pattern. Remember to keep using re.IGNORECASE to account for uppercase/lowercase variations like TAXI or taxi.
内容的提问来源于stack exchange,提问作者Anubhay Jha

