You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python BeautifulSoup爬取单网页多表格的代码适配问题求助

爬取失败核心原因

你使用的参考代码是fbref站点的专属适配逻辑:

  • fbref属于Sports Reference旗下站点,核心统计表格默认包裹在HTML注释中,必须先提取注释内容才能正常解析
  • 代码中硬编码的选择器#all_stats_standard仅对应fbref站点的英超标准统计表格容器

替换URL后失效属于正常情况:两个目标站点的页面结构、表格存储位置、反爬规则均与fbref不同,且原代码未携带浏览器请求头,直接发起请求大概率会被站点反爬策略拦截,返回无有效表格的页面内容。

适配代码(可直接在JupyterLab运行)

不需要额外安装新依赖,沿用你已有的requests、beautifulsoup4库即可。
注意:所有请求必须携带浏览器请求头,绕过基础反爬拦截

import requests
import time
from bs4 import BeautifulSoup, Comment

# 通用请求头
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/125.0.0.0 Safari/537.36"
}

适配站点1:sports-reference 球员页

该站点和fbref同属Sports Reference体系,表格同样存储在HTML注释中,仅外层容器ID不同:

  • 场均数据容器ID:#all_per_game
  • 累计数据容器ID:#all_totals
url1 = "https://www.sports-reference.com/cbb/players/jaden-ivey-1.html"
resp1 = requests.get(url1, headers=headers)
resp1.raise_for_status() # 请求失败直接抛错,方便排查
soup1 = BeautifulSoup(resp1.content, "html.parser")

# 解析场均数据表格
per_game_table = BeautifulSoup(
    soup1.select_one("#all_per_game").find_next(text=lambda x: isinstance(x, Comment)),
    "html.parser"
)

# 解析累计数据表格
total_table = BeautifulSoup(
    soup1.select_one("#all_totals").find_next(text=lambda x: isinstance(x, Comment)),
    "html.parser"
)

# 测试打印场均数据前5行
print("=== sports-reference 场均数据 ===")
for tr in per_game_table.select("tr:has(td)")[:5]:
    tds = [td.get_text(strip=True) for td in tr.select("td")]
    if len(tds) >= 6: # 跳过表头、合计等空行避免报错
        print("{:<12}{:<8}{:<8}".format(tds[0], tds[3], tds[5]))

适配站点2:realgm 球员页

该站点的统计表格直接渲染在HTML源码中,不需要解析注释,直接匹配对应table标签即可:

  • 场均数据为页面中第2个class为tablesaw的表格
  • 累计数据为页面中第3个class为tablesaw的表格
time.sleep(2) # 间隔2秒请求,避免被限流
url2 = "https://basketball.realgm.com/player/Jaden-Ivey/Summary/148740"
resp2 = requests.get(url2, headers=headers)
resp2.raise_for_status()
soup2 = BeautifulSoup(resp2.content, "html.parser")

all_tables = soup2.select("table.tablesaw")
per_game_table_realgm = all_tables[1] # 索引从0开始计数
total_table_realgm = all_tables[2]

# 测试打印场均数据前5行
print("\n=== realgm 场均数据 ===")
for tr in per_game_table_realgm.select("tr:has(td)")[:5]:
    tds = [td.get_text(strip=True) for td in tr.select("td")]
    if len(tds) >= 6:
        print("{:<12}{:<8}{:<8}".format(tds[0], tds[2], tds[4]))
新手调整提示
  • 需要提取其他列内容时,可以先打印整行tds列表查看每列对应的索引,再修改tds[索引]的取值即可
  • 如果运行时返回429状态码,说明触发站点限流,调大time.sleep()的间隔时间再重试即可
  • 若需要将表格导出为CSV/Excel格式,可以安装pandas库,用pd.read_html(str(表格变量))直接将解析后的表格转为DataFrame,不需要手动逐行拼接内容。

内容的提问来源于stack exchange,提问作者TNieland

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.28 17:57:35