如何使用正则表达式合并分组数字提取完整10位手机号码?
Absolutely, you can handle this directly with regex by first matching the possible formatted phone number patterns, then removing any spaces from the matches to get the continuous 10-digit string. Here's how to do it:
Step-by-Step Solution
First, define a regex pattern that accounts for all valid formats you mentioned: continuous 10 digits, or digits with a single space after the 2nd or 5th digit. Then clean up the matches by stripping spaces.
import re mystr = '(R) 98198 38466 (some Text) 9702977470' # Regex pattern to match valid 10-digit phone number formats phone_pattern = r'\b(?:\d{2} \d{8}|\d{5} \d{5}|\d{10})\b' # Find all matching formatted numbers formatted_numbers = re.findall(phone_pattern, mystr) # Remove spaces from each match to get continuous 10-digit numbers clean_numbers = [num.replace(' ', '') for num in formatted_numbers] print(clean_numbers) # Output: ['9819838466', '9702977470']
How the Regex Works
Let's break down the pattern \b(?:\d{2} \d{8}|\d{5} \d{5}|\d{10})\b:
\b: Word boundary to ensure we match whole numbers (not parts of longer digit sequences).(?:...): Non-capturing group to group our three valid format alternatives without creating unnecessary capture groups.\d{2} \d{8}: Matches numbers with a space after the 2nd digit (e.g.,98 12345678).\d{5} \d{5}: Matches numbers with a space after the 5th digit (e.g.,98198 38466).\d{10}: Matches continuous 10-digit numbers (e.g.,9702977470).
Why Your Original Approach Failed
Using re.findall('\d+', mystr) splits the string at any non-digit character (like the space in 98198 38466), which breaks the single phone number into two separate digit strings. The regex above targets the complete phone number in its valid formatted forms, so we can easily merge the parts by removing the space.
Edge Case Notes
- If your input might have multiple spaces instead of a single one, replace the space in the pattern with
\s+(matches one or more whitespace characters):r'\b(?:\d{2}\s+\d{8}|\d{5}\s+\d{5}|\d{10})\b'. - The
\bword boundary works well for most cases where phone numbers are surrounded by non-digit characters (like parentheses, spaces, or text). If you need to handle numbers adjacent to other characters (e.g.,abc9819838466def), you might need to adjust to lookarounds like(?<!\d)and(?!\d)instead of\b.
内容的提问来源于stack exchange,提问作者shantanuo

