Python正则提取澳大利亚手机号:解决同一子串漏匹配问题
Hey there! Let's work through why your regex is missing valid numbers and throwing false positives when parsing that web page text.
The Root of the Problem
Your current approach first strips all non-alphanumeric characters with re.sub(r'\W+', '', sample_text), which causes two big issues:
- False Positives: Unrelated numbers (like the NSW2007 postcode) get glued to adjacent phone number digits, creating fake 10-digit sequences that match your regex (like
0761028064from2007+61028064). - Missing Valid Numbers: When you flatten text like
+61 (0)2 8064 7043into610280647043, your regex only looks for612\d{8}(international format without the leading 0 in the state code) — it doesn't account for6102\d{8}(the version with the retained 0 from(0)2), so the valid number gets missed entirely.
A Better Approach: Match Patterns, Don't Flatten First
Instead of merging all digits together, we'll write a regex that matches the actual structure of Australian phone numbers (including allowed separators like spaces, brackets, and hyphens). This way, we avoid gluing unrelated digits and catch all valid formats.
Solution Code
Here's an updated regex and workflow that handles all your listed formats, plus cleans up the results consistently:
import re from selenium.webdriver.common.by import By def extract_australian_phones(text): # Regex pattern covering all valid Australian phone number formats phone_pattern = r''' # International formats: +61 (0)x xxxx xxxx, +61 x xxxx xxxx, +61 0x xxxx xxxx (?:\+61[-. ]?(?:\(0\))?[2378][-. ]?\d{4}[-. ]?\d{4}) | # Local formats: 0x xxxx xxxx, 0x-xxxx-xxxx, 0xxxxxxxxx (?:0[2378][-. ]?\d{4}[-. ]?\d{4}) ''' # Find all matches, ignoring whitespace in the pattern (re.VERBOSE) matches = re.findall(phone_pattern, text, re.VERBOSE) # Filter out empty strings (from alternation groups that didn't match) valid_matches = [match for match in matches if match] # Clean up matches to get consistent numeric formats cleaned_numbers = [] for num in valid_matches: # Remove all non-digit/non-+ characters clean = re.sub(r'[^\d+]', '', num) # Convert local numbers to international format for consistency if clean.startswith('0'): clean = '+61' + clean[1:] cleaned_numbers.append(clean) return cleaned_numbers # Example usage with your web page text sample_text = driver.find_element(By.TAG_NAME, 'body').text phones = extract_australian_phones(sample_text) print(phones)
What This Fixes
- Catches All Valid Formats: It matches local numbers (
02 1234 5678), international numbers without the leading 0 (+61 2 1234 5678), and the rare international format with(0)(+61 (0)2 1234 5678). - Reduces False Positives: By matching full number patterns instead of flat digit sequences, it won't glue postcodes or other random digits into fake phone numbers.
- Consistent Output: All numbers are converted to the international
+61xxxxxxxxxformat for uniformity, but you can adjust this to return local formats if needed.
Notes on Your Acceptable False Positives
If you still see occasional false positives (e.g., a 10-digit sequence that looks like a phone number but isn't), you can add extra checks:
- Verify the number is not part of a longer digit sequence (use lookarounds instead of word boundaries, since
+isn't a word character). - Cross-reference with known Australian number ranges (your current state code coverage for 02/03/07/08 is already solid).
内容的提问来源于stack exchange,提问作者Axle Max

