You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup与Requests爬取Fangraphs投手数据页面失效的问题咨询

使用BeautifulSoup与Requests爬取Fangraphs投手数据页面失效的问题咨询

嗨Carlos,这种情况我碰到过好多次——体育数据网站经常会悄悄调整页面结构或者升级反爬机制,咱们一步步来排查和解决:

先梳理最可能的两个失效原因

  1. 反爬拦截:请求被识别为非浏览器请求
    Requests库默认的请求头没有浏览器标识,很多网站会直接拒绝这类请求,返回空白页或者验证页面,导致你拿不到真实的表格数据。

  2. 页面结构变更:表格选择器失效
    Fangraphs大概率更新了页面的HTML结构,你之前依赖的rgMasterTable类名可能已经被修改或移除,导致代码找不到目标表格。

解决步骤和修改后的代码

方案1:优化原有BeautifulSoup代码

咱们先处理反爬问题,再调整表格和行的定位逻辑,让代码更鲁棒:

import pandas as pd
import requests
from datetime import date, timedelta
from bs4 import BeautifulSoup
import lxml
import numpy as np

def parse_array_from_fangraphs_html(start_date, end_date, URL_1):
    """
    Take a HTML stats page from fangraphs and parse it out to a dataframe.
    """
    # 添加模拟浏览器的请求头,避免被反爬拦截
    headers = {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
    }
    
    # 带请求头发起请求,请求失败直接抛出错误提示
    response = requests.get(URL_1, headers=headers)
    response.raise_for_status()
    soup = BeautifulSoup(response.text, "lxml")

    # 不再依赖固定类名,通过表头特征定位目标表格
    target_table = None
    for table in soup.find_all("table"):
        thead = table.find("thead")
        if thead:
            headers_list = [th.text.strip() for th in thead.find_all("th")]
            # 投手数据表格肯定包含Name、IP这些核心字段,用来筛选
            if "Name" in headers_list and "IP" in headers_list:
                target_table = table
                break
    
    if not target_table:
        raise ValueError("找不到目标数据表格,页面结构可能再次变更")
    
    # 提取表头
    headers = [th.text.strip() for th in target_table.find("thead").find_all("th")]
    
    # 提取数据行:Fangraphs的表格行通常用rgRow/rgAltRow类区分奇偶行
    rows = []
    rows_html = target_table.find_all("tr", class_=["rgRow", "rgAltRow"])
    for row in rows_html:
        row_data = [cell.text.strip() for cell in row.find_all("td")]
        rows.append(row_data)
    
    return pd.DataFrame(rows, columns=headers)

sdate = '2022-01-01'
enddate = date.today().strftime("%Y-%m-%d")
PITCHERS = "https://www.fangraphs.com/leaders/major-league?pos=all&stats=pit&lg=all&qual=y&type=36&season=2023&month=0&season1=2023&ind=0"
wRC1 = parse_array_from_fangraphs_html(sdate, enddate, PITCHERS)
print(wRC1.head())

方案2:用Pandas直接解析表格(更简洁)

Pandas的read_html方法可以自动提取页面中的所有表格,省去手动解析的麻烦,只要请求头正确就行:

import pandas as pd
import requests
from datetime import date

headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
}

sdate = '2022-01-01'
enddate = date.today().strftime("%Y-%m-%d")
PITCHERS = "https://www.fangraphs.com/leaders/major-league?pos=all&stats=pit&lg=all&qual=y&type=36&season=2023&month=0&season1=2023&ind=0"

# 发起请求并解析页面中的所有表格
response = requests.get(PITCHERS, headers=headers)
response.raise_for_status()
all_tables = pd.read_html(response.text)

# 筛选出投手数据表格
target_df = None
for df in all_tables:
    if "Name" in df.columns and "IP" in df.columns:
        target_df = df
        break

if target_df is not None:
    print(target_df.head())
else:
    print("未找到目标表格,请检查页面结构")

后续注意事项

  • 如果之后又失效了,先打印response.text的前几百字符,看是被反爬拦截了,还是页面结构又变了。
  • 不要频繁发起请求,避免被网站封禁IP,必要时可以加个小延迟(比如time.sleep(2))。

备注:内容来源于stack exchange,提问作者Carlos Marcano

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.21 13:03:00