You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何BeautifulSoup会将换行符解析为子节点?

关于BeautifulSoup解析换行符的特殊行为解释

先看你给出的测试代码:

>>> from bs4 import BeautifulSoup
>>> html_doc1 = """<h1>test1</h1><h1>test2</h1>"""
>>> html_doc2 = """<h1>test1</h1>\n<h1>test2</h1>"""
>>> soup1 = BeautifulSoup(html_doc1, 'html.parser')
>>> soup2 = BeautifulSoup(html_doc2, 'html.parser')
>>> print(len(list(soup1.children)), len(list(soup2.children)))
2 3

出现这种差异的原因很直接:

  • BeautifulSoup默认会保留HTML文档中的空白文本节点,换行符\n属于空白文本的一种,当HTML里存在明确的换行时,会被解析成单独的NavigableString节点,所以soup2的子节点多了一个空文本节点。
  • 你使用的html.parser是Python标准库自带的解析器,它的特性就是保留这类空白节点,不会自动忽略。

如果想让空白节点被过滤掉,有两种常用方式:

  1. 手动过滤子节点,剔除空文本:
# 过滤空白节点
filtered_soup2_children = [child for child in soup2.children if not (isinstance(child, str) and child.strip() == '')]
print(len(filtered_soup2_children))  # 输出 2
  1. 换用lxml解析器,它默认会忽略多余的空白节点(需要先通过pip install lxml安装):
soup2 = BeautifulSoup(html_doc2, 'lxml')
print(len(list(soup2.children)))  # 输出 2

内容的提问来源于stack exchange,提问作者chnlyi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 19:50:57