如何识别正则表达式中点号匹配的字符串对应字符索引
解决方案
要找出正则表达式中通配符.匹配的字符索引,可以通过以下两种方法实现:
方法一:使用标准库re(无需第三方依赖)
核心思路是:遍历字符串的每个字符,将其替换为一个原正则中其他结构无法匹配的特殊字符(比如\0),如果替换后正则仍能匹配字符串,说明该字符是被.匹配的。
import re import pandas as pd def get_dot_indices(regex_pattern, matched_str): indices = [] # 先确认原字符串确实匹配正则 if not re.fullmatch(regex_pattern, matched_str): return [] # 遍历每个字符索引 for idx in range(len(matched_str)): # 将当前字符替换为特殊字符'\0' modified_str = matched_str[:idx] + '\x00' + matched_str[idx+1:] # 检查替换后是否仍匹配正则 if re.fullmatch(regex_pattern, modified_str): indices.append(idx) return indices # 初始化DataFrame data = pd.DataFrame({ 'Name': ['A', 'B', 'C'], 'pattern_id': ['A_1', 'B_1', 'C_1'], 'regex': ['[LIVMP].{0,2}(T)P..([ST])', '[AVP].[ST][ST][ST]', '^M{0,1}([ED]).'], 'matched_string': ['PSTPLVT', 'PVTTT', 'MEG'] }) df = pd.DataFrame(data) # 新增列 df['char_indexes'] = df.apply(lambda row: get_dot_indices(row['regex'], row['matched_string']), axis=1) print(df)
运行后输出与期望结果一致:
Name pattern_id regex matched_string char_indexes 0 A A_1 [LIVMP].{0,2}(T)P..([ST]) PSTPLVT [1, 4, 5] 1 B B_1 [AVP].[ST][ST][ST] PVTTT [1] 2 C C_1 ^M{0,1}([ED]). MEG [2]
方法二:使用regex第三方库(效率更高)
如果处理大量数据,推荐使用regex库(需先通过pip install regex安装),它支持获取重复捕获组的所有匹配位置,能更高效地追踪.的匹配索引。
import regex import pandas as pd def get_dot_indices(regex_pattern, matched_str): # 替换正则中不在字符类[]内的.为捕获组(.?) # 避免替换字符类中的.(比如[.]是匹配点本身,不是通配符) modified_regex = regex.sub(r'(?<!\[)(?:[^[\]]*\])*[^[\]]*\K\.', r'(.?)', regex_pattern) match_obj = regex.fullmatch(modified_regex, matched_str) if not match_obj: return [] # 收集所有捕获组的起始索引 indices = [] for group_num in range(1, match_obj.lastindex + 1): # 获取该捕获组所有的匹配区间 for start, _ in match_obj.spans(group_num): indices.append(start) # 确保索引按顺序排列 indices.sort() return indices # 初始化DataFrame并新增列 data = pd.DataFrame({ 'Name': ['A', 'B', 'C'], 'pattern_id': ['A_1', 'B_1', 'C_1'], 'regex': ['[LIVMP].{0,2}(T)P..([ST])', '[AVP].[ST][ST][ST]', '^M{0,1}([ED]).'], 'matched_string': ['PSTPLVT', 'PVTTT', 'MEG'] }) df = pd.DataFrame(data) df['char_indexes'] = df.apply(lambda row: get_dot_indices(row['regex'], row['matched_string']), axis=1) print(df)
此方法通过将每个.转换为捕获组,利用regex库的spans()方法获取所有匹配位置,效率比标准库方法更高。
内容的提问来源于stack exchange,提问作者Bluegirl FK
相关产品推荐
相关产品推荐

