You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何替换正则匹配的特定模式首尾字符,修复hocr转PDF的XML错误

HOCR转PDF时XML转义问题解决

问题场景

我有一份Google Document AI生成的长HOCR格式XML文本,示例片段如下:

<span class='ocrx_word' id='word_1_21_0_1_0' title='bbox 409 912 417 927'>&lt;</span><span class='ocrx_word' id='word_1_21_0_1_1' title='bbox 416 911 446 925'>&lt;forest&gt;...

将其转换为可搜索PDF时,PDF库会把<forest>这类被<>包裹的文本当成损坏的XML元素报错,需要把这类内容替换为XML转义形式(比如<forest>→&lt;forest&gt;)。

已经用正则表达式(?!<(div|span|/span).*>)(<.*>)匹配到目标内容(排除合法的<span>、</span>标签),现在需要仅修改匹配内容的首尾字符,保留中间文本不变。

解决方法

1. 优化正则捕获组

调整正则表达式,把匹配内容拆分为开头的<、中间文本、**结尾的>**三个捕获组,确保精准定位需要修改的部分:

(?!<(div|span|/span)\b.*>)(<)([^>]+)(>)
  • (?!<(div|span|/span)\b.*>):负向预查,跳过合法的<div>、<span>、</span>标签
  • (<):捕获组1,匹配目标内容的开头<
  • ([^>]+):捕获组2,匹配<和>之间的所有文本
  • (>):捕获组3,匹配目标内容的结尾>

2. 替换规则

使用捕获组引用,将捕获组1替换为&lt;,捕获组3替换为&gt;,保留捕获组2的原始文本:

&lt;$2&gt;

3. 工具实现示例

Python代码

import re

# 示例HOCR内容
hocr_content = """<span class='ocrx_word' id='word_1_21_0_1_0' title='bbox 409 912 417 927'>&lt;</span><span class='ocrx_word' id='word_1_21_0_1_1' title='bbox 416 911 446 925'>&lt;forest&gt;...</span>"""
# 正则模式
pattern = r"(?!<(div|span|/span)\b.*>)(<)([^>]+)(>)"
# 执行替换
processed_content = re.sub(pattern, r"&lt;\2&gt;", hocr_content)
print(processed_content)

Linux/macOS sed命令

sed -E 's/(?!<(div|span|\/span)\b.*>)(<)([^>]+)(>)/\&lt;\2\&gt;/g' input.hocr > output.hocr

内容的提问来源于stack exchange,提问作者kang

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.10 12:05:33