You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python正则提取澳大利亚手机号:解决同一子串漏匹配问题

Fixing Australian Phone Number Extraction with Regex in Python

Hey there! Let's work through why your regex is missing valid numbers and throwing false positives when parsing that web page text.

The Root of the Problem

Your current approach first strips all non-alphanumeric characters with re.sub(r'\W+', '', sample_text), which causes two big issues:

  1. False Positives: Unrelated numbers (like the NSW2007 postcode) get glued to adjacent phone number digits, creating fake 10-digit sequences that match your regex (like 0761028064 from 2007 + 61028064).
  2. Missing Valid Numbers: When you flatten text like +61 (0)2 8064 7043 into 610280647043, your regex only looks for 612\d{8} (international format without the leading 0 in the state code) — it doesn't account for 6102\d{8} (the version with the retained 0 from (0)2), so the valid number gets missed entirely.

A Better Approach: Match Patterns, Don't Flatten First

Instead of merging all digits together, we'll write a regex that matches the actual structure of Australian phone numbers (including allowed separators like spaces, brackets, and hyphens). This way, we avoid gluing unrelated digits and catch all valid formats.

Solution Code

Here's an updated regex and workflow that handles all your listed formats, plus cleans up the results consistently:

import re
from selenium.webdriver.common.by import By

def extract_australian_phones(text):
    # Regex pattern covering all valid Australian phone number formats
    phone_pattern = r'''
        # International formats: +61 (0)x xxxx xxxx, +61 x xxxx xxxx, +61 0x xxxx xxxx
        (?:\+61[-. ]?(?:\(0\))?[2378][-. ]?\d{4}[-. ]?\d{4})
        |
        # Local formats: 0x xxxx xxxx, 0x-xxxx-xxxx, 0xxxxxxxxx
        (?:0[2378][-. ]?\d{4}[-. ]?\d{4})
    '''
    # Find all matches, ignoring whitespace in the pattern (re.VERBOSE)
    matches = re.findall(phone_pattern, text, re.VERBOSE)
    # Filter out empty strings (from alternation groups that didn't match)
    valid_matches = [match for match in matches if match]
    
    # Clean up matches to get consistent numeric formats
    cleaned_numbers = []
    for num in valid_matches:
        # Remove all non-digit/non-+ characters
        clean = re.sub(r'[^\d+]', '', num)
        # Convert local numbers to international format for consistency
        if clean.startswith('0'):
            clean = '+61' + clean[1:]
        cleaned_numbers.append(clean)
    
    return cleaned_numbers

# Example usage with your web page text
sample_text = driver.find_element(By.TAG_NAME, 'body').text
phones = extract_australian_phones(sample_text)
print(phones)

What This Fixes

  • Catches All Valid Formats: It matches local numbers (02 1234 5678), international numbers without the leading 0 (+61 2 1234 5678), and the rare international format with (0) (+61 (0)2 1234 5678).
  • Reduces False Positives: By matching full number patterns instead of flat digit sequences, it won't glue postcodes or other random digits into fake phone numbers.
  • Consistent Output: All numbers are converted to the international +61xxxxxxxxx format for uniformity, but you can adjust this to return local formats if needed.

Notes on Your Acceptable False Positives

If you still see occasional false positives (e.g., a 10-digit sequence that looks like a phone number but isn't), you can add extra checks:

  • Verify the number is not part of a longer digit sequence (use lookarounds instead of word boundaries, since + isn't a word character).
  • Cross-reference with known Australian number ranges (your current state code coverage for 02/03/07/08 is already solid).

内容的提问来源于stack exchange,提问作者Axle Max

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 04:47:30