You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Fangraphs网站改版后Python爬虫无法爬取新排行榜求助

爬取Fangraphs新版排行榜表格的问题解决建议

背景

此前我用以下Python脚本通过requests爬取Fangraphs旧版兼容链接的表格完全正常:

import requests
import datetime
from datetime import date, timedelta
from bs4 import BeautifulSoup
import pandas as pd
import lxml
import openpyxl

def parse_array_from_fangraphs_html(start_date,end_date):
    """
    Take a HTML stats page from fangraphs and parse it out to a dataframe.
    """
    # parse input
  
    #PITCHERS_URL = "https://www.fangraphs.com/leaders/major-league?pos=all&stats=pit&lg=all&qual=0&type=c%2C13%2C7%2C8%2C120%2C121%2C331%2C105%2C111%2C24%2C19%2C14%2C329%2C324%2C45%2C122%2C6%2C42%2C43%2C328%2C330%2C322%2C323%2C326%2C332%2C31%2C30%2C29&season=2021&month=1000&season1=2015&ind=0&team=&rost=&age=&filter=&players=&startdate={}&enddate={}&page=1_5000&pagenum=1&pageitems=2000000000".format(start_date, end_date)
    PITCHERS_URL = "https://www.fangraphs.com/leaders-legacy.aspx?pos=all&stats=pit&lg=all&qual=0&type=c,13,7,8,120,121,331,105,111,24,19,14,329,324,45,122,6,42,43,328,330,322,323,326,332,31,30,29&season=2021&month=1000&season1=2015&ind=0&team=&rost=&age=&filter=&players=&startdate={}&enddate={}&page=1_2000".format(start_date, end_date)   
    # request the data
    pitchers_html = requests.get(PITCHERS_URL).text
    soup = BeautifulSoup(pitchers_html, "lxml")
    table = soup.find("table", {"class": "rgMasterTable"})
    
    # get headers
    headers_html = table.find("thead").find_all("th")
    headers = []
    for header in headers_html:
        headers.append(header.text)

    # get rows
    rows = []
    rows_html = table.find("tbody").find_all("tr")
    for row in rows_html:
        row_data = []
        for cell in row.find_all("td"):
            row_data.append(cell.text)
        rows.append(row_data)
    
    return pd.DataFrame(rows, columns = headers)

def calc_speX(df, IP_limit):
    # A bunch of calculations
    
    return speX

sdate = '2023-03-30'
enddate = '2023-08-17'
IP = 0 
date_format = "%Y-%m-%d"
start_date = datetime.datetime.strptime(sdate, date_format)
end_date = datetime.datetime.strptime(enddate, date_format)

daily = input("Do you want the daily values? ")
writer = pd.ExcelWriter('speX-daily.xlsx', engine='openpyxl') 

if daily.lower()=="y":
    for single_date in pd.date_range(start=start_date, end=end_date):
        date_str = single_date.strftime(date_format)
        speX = parse_array_from_fangraphs_html(date_str, date_str)
        result = calc_speX(speX, IP)
        result.to_excel(writer, sheet_name=date_str)

writer.save()
writer.close()

但网站改版后的新版排行榜链接无法用requests直接爬取,尝试用Selenium渲染动态内容也失败,对应的Selenium代码如下:

chrome_options = Options() 
chrome_options.add_argument("--headless") 
chromedriver_path = '/chd/chromedriver.exe' 
driver = webdriver.Chrome(executable_path=chromedriver_path, options=chrome_options) 
driver.get(url) 
time.sleep(5) 
html = driver.page_source 
driver.quit() 
soup = BeautifulSoup(html, "lxml") 
table = soup.find("table", {"class": "rgMasterTable"})

运行后获取到的table为空。

可能的原因

  1. 新版页面的表格class名已变更,不再使用rgMasterTable
  2. 新版页面启用了反爬机制,检测到headless浏览器并限制内容加载
  3. 表格数据通过异步API加载,未直接渲染在初始HTML中

解决建议

1. 确认新版页面的表格选择器

打开新版页面的开发者工具(F12),在Elements面板中定位表格元素:

  • 查找表格对应的真实class、id或其他属性,替换Soup中的查找条件
  • 若页面用div模拟表格结构,需调整解析逻辑

2. 优化Selenium配置绕反爬

更新Selenium参数,模拟真实浏览器环境:

from selenium.webdriver.chrome.options import Options
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

chrome_options = Options()
# 使用新版headless模式,更接近真实浏览器
chrome_options.add_argument("--headless=new")
# 禁用自动化检测特征
chrome_options.add_argument("--disable-blink-features=AutomationControlled")
# 设置真实用户代理
chrome_options.add_argument("--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/117.0.0.0 Safari/537.36")
chrome_options.add_experimental_option("excludeSwitches", ["enable-automation"])
chrome_options.add_experimental_option('useAutomationExtension', False)

chromedriver_path = '/chd/chromedriver.exe'
driver = webdriver.Chrome(executable_path=chromedriver_path, options=chrome_options)
driver.get(url)

# 改用显式等待,直到表格加载完成(替换time.sleep)
try:
    # 替换"new-table-class"为实际表格class
    wait = WebDriverWait(driver, 15)
    table_element = wait.until(EC.presence_of_element_located((By.CLASS_NAME, "new-table-class")))
    html = driver.page_source
finally:
    driver.quit()

soup = BeautifulSoup(html, "lxml")
table = soup.find("table", {"class": "new-table-class"})

3. 抓包获取API接口

通过开发者工具Network面板抓包:

  1. 刷新新版页面,筛选XHR/Fetch请求
  2. 找到返回表格数据的JSON接口
  3. 复制接口的请求URL、headers和参数,用requests直接调用获取数据,效率远高于Selenium
  4. 注意保留必要的请求头(如Referer、Cookie等)以通过验证

4. 临时继续使用旧版链接

若旧版兼容链接仍可用,可暂时继续使用,定期检查链接有效性即可

内容的提问来源于stack exchange,提问作者Carlos Marcano

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 18:47:32