You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python实现HTML标签替换并排除特定<br><br>行的问题

解决HTML标签替换的Python代码修正问题

问题背景

需要处理test.html文件:

  • 定位所有以<p class="text_obisnuit">或<p class="text_obisnuit2">开头的行
  • 这类行原本应该用</p>闭合,但错误地使用了<br>,需要将行尾的<br>替换为</p>
  • 例外情况:仅包含<br><br>的行需保持原样,不能被错误替换

示例内容

原始HTML片段

<p class="text_obisnuit">这是一段文本<br>
<p class="text_obisnuit2">另一段文本<br>
<br><br>
<p class="text_obisnuit">第三段<br>

现有错误代码

import re

with open('test.html', 'r', encoding='utf-8') as f:
    content = f.readlines()

modified = []
for line in content:
    # 匹配以指定p标签开头的行
    if re.match(r'^<p class="text_obisnuit(2)?">', line):
        # 替换<br>为</p>
        line = re.sub(r'<br>\s*$', '</p>\n', line)
    modified.append(line)

with open('test_modified.html', 'w', encoding='utf-8') as f:
    f.writelines(modified)

错误输出

<p class="text_obisnuit">这是一段文本</p>
<p class="text_obisnuit2">另一段文本</p>
<p class="text_obisnuit"></p>
<p class="text_obisnuit">第三段</p>

期望输出

<p class="text_obisnuit">这是一段文本</p>
<p class="text_obisnuit2">另一段文本</p>
<br><br>
<p class="text_obisnuit">第三段</p>

修正后的代码

import re

with open('test.html', 'r', encoding='utf-8') as f:
    content = f.readlines()

modified = []
for line in content:
    stripped_line = line.strip()
    # 跳过仅含<br><br>的行
    if stripped_line == '<br><br>':
        modified.append(line)
        continue
    # 仅处理以指定p标签开头的行
    if re.match(r'^<p class="text_obisnuit(2)?">', line):
        # 替换行尾的<br>为</p>
        line = re.sub(r'<br>\s*$', '</p>\n', line)
    modified.append(line)

with open('test_modified.html', 'w', encoding='utf-8') as f:
    f.writelines(modified)

修正逻辑说明

  1. 新增跳过判断:先对每行做去除首尾空白的处理,判断是否等于<br><br>,如果是则直接保留原行,不进入后续替换逻辑
  2. 保留原有匹配逻辑:仅对以指定p标签开头的行执行<br>到</p>的替换,确保无关行不受影响

这样就能避免把仅含<br><br>的行错误替换成空的p标签,同时完成目标标签的修正。

内容的提问来源于stack exchange,提问作者Just Me

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 17:32:43