You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用lxml的etree.HTML解析网页时如何保留原有HTML标签结构?

问题解决方案

根因说明

etree.HTML()默认启用HTML容错修复机制,当待解析页面存在未闭合标签、嵌套错误等不规范语法时,lxml会自动补全标签、调整节点顺序,最终导致原始<article>标签的范围被错误扩大,出现提取内容超出预期的情况。

解决方案

方案1:自定义解析器保留原始结构

如果需要严格保留原始HTML的节点结构,不允许lxml自动修复,可以自定义HTMLParser关闭容错模式:

import requests
from lxml import etree
res = requests.get('https://www.nepalitimes.com/here-now/a-short-walk-up-the-panjshir/')
# 自定义解析器,关闭自动容错修复,指定页面编码避免乱码
parser = etree.HTMLParser(recover=False, encoding=res.encoding)
html = etree.fromstring(res.text, parser=parser)
# 此时解析后的结构和原始页面完全一致,可直接使用原有xpath逻辑
article_node = html.xpath('//article')[1]
print(etree.tostring(article_node, encoding='unicode'))

注意:如果原始页面HTML语法错误过多,关闭容错模式会直接抛出解析异常,这种场景推荐使用方案2。

方案2:精准xpath定位绕开结构修改影响

如果只需要正确提取正文,不需要纠结解析器是否修改结构,可以通过标签属性做精准筛选,不需要依赖节点顺序:
该页面正文<article>标签带有class="post"专属属性,直接通过属性筛选即可准确定位到正文节点:

import requests
from lxml import etree
res = requests.get('https://www.nepalitimes.com/here-now/a-short-walk-up-the-panjshir/')
html = etree.HTML(res.text)
# 带属性筛选的xpath,不受节点结构调整影响
article_node = html.xpath('//article[contains(@class, "post")]')[0]
print(etree.tostring(article_node, encoding='unicode'))

方案3:改用BeautifulSoup解析

如果lxml的结构调整问题始终无法解决,可以换用BeautifulSoup解析器,它对不规范HTML的容错处理逻辑更贴合浏览器渲染逻辑,结构保留更接近原始页面:

import requests
from bs4 import BeautifulSoup
res = requests.get('https://www.nepalitimes.com/here-now/a-short-walk-up-the-panjshir/')
soup = BeautifulSoup(res.text, 'lxml')
article_content = soup.find('article', class_='post')
print(article_content.get_text())

内容的提问来源于stack exchange,提问作者old_wang

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 15:12:00