如何替换正则匹配的特定模式首尾字符,修复hocr转PDF的XML错误
HOCR转PDF时XML转义问题解决
问题场景
我有一份Google Document AI生成的长HOCR格式XML文本,示例片段如下:
<span class='ocrx_word' id='word_1_21_0_1_0' title='bbox 409 912 417 927'><</span><span class='ocrx_word' id='word_1_21_0_1_1' title='bbox 416 911 446 925'><forest>...
将其转换为可搜索PDF时,PDF库会把<forest>这类被<>包裹的文本当成损坏的XML元素报错,需要把这类内容替换为XML转义形式(比如<forest>→<forest>)。
已经用正则表达式(?!<(div|span|/span).*>)(<.*>)匹配到目标内容(排除合法的<span>、</span>标签),现在需要仅修改匹配内容的首尾字符,保留中间文本不变。
解决方法
1. 优化正则捕获组
调整正则表达式,把匹配内容拆分为开头的<、中间文本、**结尾的>**三个捕获组,确保精准定位需要修改的部分:
(?!<(div|span|/span)\b.*>)(<)([^>]+)(>)
(?!<(div|span|/span)\b.*>):负向预查,跳过合法的<div>、<span>、</span>标签(<):捕获组1,匹配目标内容的开头<([^>]+):捕获组2,匹配<和>之间的所有文本(>):捕获组3,匹配目标内容的结尾>
2. 替换规则
使用捕获组引用,将捕获组1替换为<,捕获组3替换为>,保留捕获组2的原始文本:
<$2>
3. 工具实现示例
Python代码
import re # 示例HOCR内容 hocr_content = """<span class='ocrx_word' id='word_1_21_0_1_0' title='bbox 409 912 417 927'><</span><span class='ocrx_word' id='word_1_21_0_1_1' title='bbox 416 911 446 925'><forest>...</span>""" # 正则模式 pattern = r"(?!<(div|span|/span)\b.*>)(<)([^>]+)(>)" # 执行替换 processed_content = re.sub(pattern, r"<\2>", hocr_content) print(processed_content)
Linux/macOS sed命令
sed -E 's/(?!<(div|span|\/span)\b.*>)(<)([^>]+)(>)/\<\2\>/g' input.hocr > output.hocr
内容的提问来源于stack exchange,提问作者kang
相关产品推荐
相关产品推荐

