You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何识别正则表达式中点号匹配的字符串对应字符索引

解决方案

要找出正则表达式中通配符.匹配的字符索引,可以通过以下两种方法实现:

方法一:使用标准库re(无需第三方依赖)

核心思路是:遍历字符串的每个字符,将其替换为一个原正则中其他结构无法匹配的特殊字符(比如\0),如果替换后正则仍能匹配字符串,说明该字符是被.匹配的。

import re
import pandas as pd

def get_dot_indices(regex_pattern, matched_str):
    indices = []
    # 先确认原字符串确实匹配正则
    if not re.fullmatch(regex_pattern, matched_str):
        return []
    # 遍历每个字符索引
    for idx in range(len(matched_str)):
        # 将当前字符替换为特殊字符'\0'
        modified_str = matched_str[:idx] + '\x00' + matched_str[idx+1:]
        # 检查替换后是否仍匹配正则
        if re.fullmatch(regex_pattern, modified_str):
            indices.append(idx)
    return indices

# 初始化DataFrame
data = pd.DataFrame({
    'Name': ['A', 'B', 'C'], 
    'pattern_id': ['A_1', 'B_1', 'C_1'], 
    'regex': ['[LIVMP].{0,2}(T)P..([ST])', '[AVP].[ST][ST][ST]', '^M{0,1}([ED]).'], 
    'matched_string': ['PSTPLVT', 'PVTTT', 'MEG']
})
df = pd.DataFrame(data)

# 新增列
df['char_indexes'] = df.apply(lambda row: get_dot_indices(row['regex'], row['matched_string']), axis=1)

print(df)

运行后输出与期望结果一致:

Name pattern_id                        regex matched_string char_indexes
0    A        A_1  [LIVMP].{0,2}(T)P..([ST])        PSTPLVT  [1, 4, 5]
1    B        B_1         [AVP].[ST][ST][ST]          PVTTT        [1]
2    C        C_1             ^M{0,1}([ED]).            MEG        [2]

方法二:使用regex第三方库(效率更高)

如果处理大量数据,推荐使用regex库(需先通过pip install regex安装),它支持获取重复捕获组的所有匹配位置,能更高效地追踪.的匹配索引。

import regex
import pandas as pd

def get_dot_indices(regex_pattern, matched_str):
    # 替换正则中不在字符类[]内的.为捕获组(.?)
    # 避免替换字符类中的.(比如[.]是匹配点本身,不是通配符)
    modified_regex = regex.sub(r'(?<!\[)(?:[^[\]]*\])*[^[\]]*\K\.', r'(.?)', regex_pattern)
    match_obj = regex.fullmatch(modified_regex, matched_str)
    if not match_obj:
        return []
    # 收集所有捕获组的起始索引
    indices = []
    for group_num in range(1, match_obj.lastindex + 1):
        # 获取该捕获组所有的匹配区间
        for start, _ in match_obj.spans(group_num):
            indices.append(start)
    # 确保索引按顺序排列
    indices.sort()
    return indices

# 初始化DataFrame并新增列
data = pd.DataFrame({
    'Name': ['A', 'B', 'C'], 
    'pattern_id': ['A_1', 'B_1', 'C_1'], 
    'regex': ['[LIVMP].{0,2}(T)P..([ST])', '[AVP].[ST][ST][ST]', '^M{0,1}([ED]).'], 
    'matched_string': ['PSTPLVT', 'PVTTT', 'MEG']
})
df = pd.DataFrame(data)
df['char_indexes'] = df.apply(lambda row: get_dot_indices(row['regex'], row['matched_string']), axis=1)

print(df)

此方法通过将每个.转换为捕获组,利用regex库的spans()方法获取所有匹配位置,效率比标准库方法更高。

内容的提问来源于stack exchange,提问作者Bluegirl FK

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 15:17:08