You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在下划线分隔的字符串中精准匹配3位与4位数字子串

解决方案

核心思路

你之前的问题出在没限定匹配的是整个子串,导致正则会从4位数字串里截出前3位。要解决这个,就得确保我们匹配的是完全由3位或4位数字组成的独立子串——也就是子串的前后要么是下划线,要么是字符串的开头/结尾。

实现代码

方法一:分割后逐个校验(直观好懂)

import re

def get_target_digits(s):
    parts = s.split('_')
    three_digit = None
    four_digit = None
    
    for part in parts:
        # 检查子串是否正好是3位数字
        if re.fullmatch(r'\d{3}', part):
            three_digit = part
        # 检查子串是否正好是4位数字
        elif re.fullmatch(r'\d{4}', part):
            four_digit = part
    
    # 只有两个目标子串都存在时才返回结果
    if three_digit and four_digit:
        return {'3位数字': three_digit, '4位数字': four_digit}
    return None

# 测试你的示例字符串
test_cases = [
    '1432_ott_457_blusp',
    '312_fooob_bork_1234',
    'broz_901_6453',
    'kkhas_1781_LET_GROK_234',
    '1781_234_kkhas'
]

for case in test_cases:
    print(f"字符串: {case}")
    print(f"匹配结果: {get_target_digits(case)}\n")

方法二:正则直接匹配(高效简洁)

用正则的边界断言,直接定位符合要求的独立子串:

import re

def get_target_digits(s):
    # 匹配前后是下划线/字符串首尾的3位数字子串
    three_match = re.search(r'(?<=_|^)\d{3}(?=_|$)', s)
    # 匹配前后是下划线/字符串首尾的4位数字子串
    four_match = re.search(r'(?<=_|^)\d{4}(?=_|$)', s)
    
    if three_match and four_match:
        return {'3位数字': three_match.group(), '4位数字': four_match.group()}
    return None

# 测试代码同上

关键说明

  • re.fullmatch:强制要求整个子串完全匹配正则,避免从长数字串里截取部分内容。
  • (?<=_|^)和(?=_|$):这两个是正向断言,用来确保我们匹配的数字子串是被下划线或字符串边界包裹的,保证它是独立的子串,不是更长内容的一部分。
  • 两种方法都会先校验是否同时存在3位和4位数字子串,符合你“先验证前提”的需求。

内容的提问来源于stack exchange,提问作者DeltaIV

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 21:05:20