You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Beautiful Soup爬取网页第三个表格失败的问题求助

解决方法

问题出在目标表格被网站放在了HTML注释中,而非直接的DOM结构里,常规的find方法无法直接定位到注释内的标签。以下是具体解决步骤:

1. 核心原因

Pro Football Reference网站会将非默认展示的表格(比如你要爬的coaching_history)包裹在HTML注释<!-- ... -->中,直接解析页面DOM时会忽略注释内容,导致无法找到目标表格。

2. 代码实现

使用Beautiful Soup的Comment类识别注释节点,提取其中的HTML内容后重新解析:

import requests
from bs4 import BeautifulSoup, Comment

# 模拟浏览器请求,避免反爬
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
}
url = "https://www.pro-football-reference.com/coaches/ReidAn0.htm"
response = requests.get(url, headers=headers)
soup = BeautifulSoup(response.text, 'html.parser')

# 定位包含目标表格的注释块
target_comment = soup.find(string=lambda text: isinstance(text, Comment) and 'coaching_history' in text)

# 解析注释内的HTML内容
comment_soup = BeautifulSoup(target_comment, 'html.parser')

# 提取目标表格
table2 = comment_soup.find('table', id="coaching_history")

# 验证并处理表格数据(示例)
if table2:
    print("成功获取目标表格")
    rows = table2.find_all('tr')
    for row in rows:
        cols = row.find_all('td')
        if cols:
            print([col.text.strip() for col in cols])
else:
    print("未找到目标表格")

3. 注意事项

  • 其他隐藏表格(如页面内的后续表格)也可以用相同方法处理,只需替换注释查找逻辑中的表格ID即可。
  • 务必添加User-Agent请求头,避免被网站识别为爬虫而拦截请求。

内容的提问来源于stack exchange,提问作者Maddie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 23:15:58