You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

正则匹配完整字符串问题:手机号正则匹配结果不符求助

Let's break down what's going wrong with your regex and fix it step by step!

What's Causing the Issues?

  1. Unwanted Capture Groups: Your regex r'((\+91|0)?\s?\d{10})' has two nested capture groups. When you use re.findall(), it returns all captured groups instead of just the full match. That's why you're getting tuples like ('+91 1234567890', '+91') instead of just the complete phone number.
  2. No Boundary Checks: Without boundary markers (like \b for word boundaries), your regex will happily match any 10-digit substring inside a longer number. That's why 1234568901112 gets chopped down to 1234568901—it's grabbing the first 10 digits it finds.
  3. Incorrect Length Logic: Your pattern treats 0 as an optional prefix followed by 10 digits, which means it will match 0123456789 (10 digits total) instead of the intended 11-digit 01234567890. The logic for 0-starting numbers needs to be explicitly defined as 11 digits.

The Fixed Regex & Solution

We need to adjust the regex to:

  • Use non-capture groups to avoid extra tuple entries
  • Add boundary checks to prevent partial matches
  • Explicitly define each valid phone number format

Here's the corrected implementation:

import re

target_text = """+91 1234567890
1234567790
01234567890
1234568901112"""

# Fixed regex pattern
phone_pattern = r'\b(?:\+91\s?\d{10}|0\d{10}|\d{10})\b'
matches = re.findall(phone_pattern, target_text)

print(matches)
# Output: ['+91 1234567890', '1234567790', '01234567890']

Let's Break Down the Fixed Pattern

  • \b: Word boundary marker—ensures we only match standalone phone numbers, not parts of longer strings.
  • (?:...): Non-capturing group—groups the different phone formats together without capturing extra data, so findall() returns just the full matches.
  • \+91\s?\d{10}: Matches numbers with the +91 country code, followed by an optional space and 10 digits.
  • 0\d{10}: Matches 11-digit numbers starting with 0 (exactly 1 digit for the prefix + 10 digits).
  • \d{10}: Matches standard 10-digit numbers with no prefix.

If You Need to Validate Entire Lines

If you're checking that a whole string is a valid phone number (not extracting from a larger text), use line boundaries instead of word boundaries:

line_validation_pattern = r'^(?:\+91\s?\d{10}|0\d{10}|\d{10})$'

That should give you exactly the matches you're expecting!

内容的提问来源于stack exchange,提问作者Mohit Motwani

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 04:40:37