Python正则匹配<a>标签内容后计算其在无标签文本中的对应索引
解决方案
核心思路
你要计算的偏移量本质是匹配项对应位置之前所有<a>和</a>标签的总长度:
- 每个
<a>占3个字符,每个</a>占4个字符 - 匹配项在原文本的起始位置减去该位置前所有标签的总长度,就是纯文本里的起始位置
- 匹配项在原文本的结束位置减去该位置前所有标签的总长度,就是纯文本里的结束位置
该方法完全基于原文本的位置计算,不需要二次搜索纯文本,从根源避免了重复内容匹配错误的问题。
完整实现代码
import re def get_annotated_positions(raw_text): # 匹配所有<a>标签包裹的内容 content_pattern = re.compile(r'(?<=<a>).*?(?=</a>)') content_matches = list(content_pattern.finditer(raw_text)) # 预存所有开、闭标签的起始位置 open_tag_pos = [m.start() for m in re.finditer(r'<a>', raw_text)] close_tag_pos = [m.start() for m in re.finditer(r'</a>', raw_text)] result = [] for match in content_matches: orig_start = match.start() orig_end = match.end() # 计算起始位置对应的偏移 open_before_start = sum(1 for s in open_tag_pos if s < orig_start) close_before_start = sum(1 for s in close_tag_pos if s < orig_start) offset_start = open_before_start * 3 + close_before_start * 4 pure_start = orig_start - offset_start # 计算结束位置对应的偏移 open_before_end = sum(1 for s in open_tag_pos if s < orig_end) close_before_end = sum(1 for s in close_tag_pos if s < orig_end) offset_end = open_before_end * 3 + close_before_end * 4 pure_end = orig_end - offset_end result.append({ "content": match.group(), "start": pure_start, "end": pure_end }) # 同步返回去标签后的纯文本用于验证 pure_text = re.sub(r'</?a>', '', raw_text) return result, pure_text
测试示例
raw_text = "<a>George</a> and his <a>friends</a>, came back <a>home</a>." positions, pure_txt = get_annotated_positions(raw_text) print("纯文本内容:", pure_txt) # 输出:纯文本内容: George and his friends, came back home. for item in positions: print(item) # 切片验证正确性 print(f"位置验证:{pure_txt[item['start']:item['end']]}")
运行后输出的位置信息完全匹配需求,切片验证可以直接确认位置正确性。
大文本性能优化
如果处理的文本体积很大,可以用二分查找替代遍历计数,大幅提升计算速度:
import bisect # 把之前的sum计数替换为bisect_left即可 open_before_start = bisect.bisect_left(open_tag_pos, orig_start) close_before_start = bisect.bisect_left(close_tag_pos, orig_start)
内容的提问来源于stack exchange,提问作者Paschalis
相关产品推荐
相关产品推荐

