使用lxml的etree.HTML解析网页时如何保留原有HTML标签结构?
问题解决方案
根因说明
etree.HTML()默认启用HTML容错修复机制,当待解析页面存在未闭合标签、嵌套错误等不规范语法时,lxml会自动补全标签、调整节点顺序,最终导致原始<article>标签的范围被错误扩大,出现提取内容超出预期的情况。
解决方案
方案1:自定义解析器保留原始结构
如果需要严格保留原始HTML的节点结构,不允许lxml自动修复,可以自定义HTMLParser关闭容错模式:
import requests from lxml import etree res = requests.get('https://www.nepalitimes.com/here-now/a-short-walk-up-the-panjshir/') # 自定义解析器,关闭自动容错修复,指定页面编码避免乱码 parser = etree.HTMLParser(recover=False, encoding=res.encoding) html = etree.fromstring(res.text, parser=parser) # 此时解析后的结构和原始页面完全一致,可直接使用原有xpath逻辑 article_node = html.xpath('//article')[1] print(etree.tostring(article_node, encoding='unicode'))
注意:如果原始页面HTML语法错误过多,关闭容错模式会直接抛出解析异常,这种场景推荐使用方案2。
方案2:精准xpath定位绕开结构修改影响
如果只需要正确提取正文,不需要纠结解析器是否修改结构,可以通过标签属性做精准筛选,不需要依赖节点顺序:
该页面正文<article>标签带有class="post"专属属性,直接通过属性筛选即可准确定位到正文节点:
import requests from lxml import etree res = requests.get('https://www.nepalitimes.com/here-now/a-short-walk-up-the-panjshir/') html = etree.HTML(res.text) # 带属性筛选的xpath,不受节点结构调整影响 article_node = html.xpath('//article[contains(@class, "post")]')[0] print(etree.tostring(article_node, encoding='unicode'))
方案3:改用BeautifulSoup解析
如果lxml的结构调整问题始终无法解决,可以换用BeautifulSoup解析器,它对不规范HTML的容错处理逻辑更贴合浏览器渲染逻辑,结构保留更接近原始页面:
import requests from bs4 import BeautifulSoup res = requests.get('https://www.nepalitimes.com/here-now/a-short-walk-up-the-panjshir/') soup = BeautifulSoup(res.text, 'lxml') article_content = soup.find('article', class_='post') print(article_content.get_text())
内容的提问来源于stack exchange,提问作者old_wang
相关产品推荐
相关产品推荐

