为何BeautifulSoup的find_all()方法在HTML注释标签后停止解析?
问题分析与解决:BeautifulSoup解析时丢失后续元素
问题根源
你使用的Python内置html.parser解析器,对超长且包含不规范内容的HTML注释容错性较差,导致注释后的DOM元素被错误包裹进注释节点中,遍历父节点子元素时无法获取到这些内容。
解决方法
方法1:更换更健壮的解析器
改用lxml或html5lib解析器,它们对不规范HTML的处理能力更强。需要先安装对应库:
pip install lxml # 或者 pip install html5lib
修改后的代码:
import requests from bs4 import BeautifulSoup url = "https://www.baseball-reference.com/postseason/1905_WS.shtml" response = requests.get(url) response.raise_for_status() # 使用lxml解析器 soup = BeautifulSoup(response.content, "lxml") pitching = soup.find_all("div", id=lambda x: x and x.startswith("all_post_pitching_"))[0] # 遍历子元素,仅输出div节点 for div in pitching.find_all("div", recursive=False): print(div)
方法2:手动跳过注释节点获取后续元素
如果不想更换解析器,可以遍历父节点的所有内容节点,找到注释后提取后续的兄弟元素:
import requests from bs4 import BeautifulSoup, Comment url = "https://www.baseball-reference.com/postseason/1905_WS.shtml" response = requests.get(url) response.raise_for_status() soup = BeautifulSoup(response.content, "html.parser") pitching = soup.find_all("div", id=lambda x: x and x.startswith("all_post_pitching_"))[0] found_comment = False for node in pitching.contents: # 标记已找到超长注释 if isinstance(node, Comment): found_comment = True continue # 输出注释后的所有div节点 if found_comment and node.name == "div": print(node)
补充说明
原代码中直接遍历pitching的子元素时,由于html.parser解析错误,后续的<div>节点被包含在注释节点内部,因此无法被正常遍历到。更换解析器或手动处理注释节点,能有效解决这个问题。
内容的提问来源于stack exchange,提问作者Anthony
相关产品推荐
相关产品推荐

