You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python正则匹配<a>标签内容后计算其在无标签文本中的对应索引

解决方案

核心思路

你要计算的偏移量本质是匹配项对应位置之前所有<a>和</a>标签的总长度:

  • 每个<a>占3个字符,每个</a>占4个字符
  • 匹配项在原文本的起始位置减去该位置前所有标签的总长度,就是纯文本里的起始位置
  • 匹配项在原文本的结束位置减去该位置前所有标签的总长度,就是纯文本里的结束位置
    该方法完全基于原文本的位置计算,不需要二次搜索纯文本,从根源避免了重复内容匹配错误的问题。

完整实现代码

import re

def get_annotated_positions(raw_text):
    # 匹配所有<a>标签包裹的内容
    content_pattern = re.compile(r'(?<=<a>).*?(?=</a>)')
    content_matches = list(content_pattern.finditer(raw_text))
    # 预存所有开、闭标签的起始位置
    open_tag_pos = [m.start() for m in re.finditer(r'<a>', raw_text)]
    close_tag_pos = [m.start() for m in re.finditer(r'</a>', raw_text)]
    
    result = []
    for match in content_matches:
        orig_start = match.start()
        orig_end = match.end()
        
        # 计算起始位置对应的偏移
        open_before_start = sum(1 for s in open_tag_pos if s < orig_start)
        close_before_start = sum(1 for s in close_tag_pos if s < orig_start)
        offset_start = open_before_start * 3 + close_before_start * 4
        pure_start = orig_start - offset_start
        
        # 计算结束位置对应的偏移
        open_before_end = sum(1 for s in open_tag_pos if s < orig_end)
        close_before_end = sum(1 for s in close_tag_pos if s < orig_end)
        offset_end = open_before_end * 3 + close_before_end * 4
        pure_end = orig_end - offset_end
        
        result.append({
            "content": match.group(),
            "start": pure_start,
            "end": pure_end
        })
    # 同步返回去标签后的纯文本用于验证
    pure_text = re.sub(r'</?a>', '', raw_text)
    return result, pure_text

测试示例

raw_text = "<a>George</a> and his <a>friends</a>, came back <a>home</a>."
positions, pure_txt = get_annotated_positions(raw_text)

print("纯文本内容:", pure_txt)
# 输出:纯文本内容: George and his friends, came back home.
for item in positions:
    print(item)
    # 切片验证正确性
    print(f"位置验证:{pure_txt[item['start']:item['end']]}")

运行后输出的位置信息完全匹配需求,切片验证可以直接确认位置正确性。

大文本性能优化

如果处理的文本体积很大,可以用二分查找替代遍历计数,大幅提升计算速度:

import bisect

# 把之前的sum计数替换为bisect_left即可
open_before_start = bisect.bisect_left(open_tag_pos, orig_start)
close_before_start = bisect.bisect_left(close_tag_pos, orig_start)

内容的提问来源于stack exchange,提问作者Paschalis

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.02 13:06:03