You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何BeautifulSoup的find_all()方法在HTML注释标签后停止解析?

问题分析与解决:BeautifulSoup解析时丢失后续元素

问题根源

你使用的Python内置html.parser解析器,对超长且包含不规范内容的HTML注释容错性较差,导致注释后的DOM元素被错误包裹进注释节点中,遍历父节点子元素时无法获取到这些内容。

解决方法

方法1:更换更健壮的解析器

改用lxml或html5lib解析器,它们对不规范HTML的处理能力更强。需要先安装对应库:

pip install lxml
# 或者 pip install html5lib

修改后的代码:

import requests
from bs4 import BeautifulSoup

url = "https://www.baseball-reference.com/postseason/1905_WS.shtml"
response = requests.get(url)
response.raise_for_status()

# 使用lxml解析器
soup = BeautifulSoup(response.content, "lxml")

pitching = soup.find_all("div", id=lambda x: x and x.startswith("all_post_pitching_"))[0]
# 遍历子元素,仅输出div节点
for div in pitching.find_all("div", recursive=False):
    print(div)

方法2:手动跳过注释节点获取后续元素

如果不想更换解析器,可以遍历父节点的所有内容节点,找到注释后提取后续的兄弟元素:

import requests
from bs4 import BeautifulSoup, Comment

url = "https://www.baseball-reference.com/postseason/1905_WS.shtml"
response = requests.get(url)
response.raise_for_status()

soup = BeautifulSoup(response.content, "html.parser")

pitching = soup.find_all("div", id=lambda x: x and x.startswith("all_post_pitching_"))[0]
found_comment = False
for node in pitching.contents:
    # 标记已找到超长注释
    if isinstance(node, Comment):
        found_comment = True
        continue
    # 输出注释后的所有div节点
    if found_comment and node.name == "div":
        print(node)

补充说明

原代码中直接遍历pitching的子元素时,由于html.parser解析错误,后续的<div>节点被包含在注释节点内部,因此无法被正常遍历到。更换解析器或手动处理注释节点,能有效解决这个问题。

内容的提问来源于stack exchange,提问作者Anthony

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.16 14:52:12