Fangraphs网站改版后Python爬虫无法爬取新排行榜求助
爬取Fangraphs新版排行榜表格的问题解决建议
背景
此前我用以下Python脚本通过requests爬取Fangraphs旧版兼容链接的表格完全正常:
import requests import datetime from datetime import date, timedelta from bs4 import BeautifulSoup import pandas as pd import lxml import openpyxl def parse_array_from_fangraphs_html(start_date,end_date): """ Take a HTML stats page from fangraphs and parse it out to a dataframe. """ # parse input #PITCHERS_URL = "https://www.fangraphs.com/leaders/major-league?pos=all&stats=pit&lg=all&qual=0&type=c%2C13%2C7%2C8%2C120%2C121%2C331%2C105%2C111%2C24%2C19%2C14%2C329%2C324%2C45%2C122%2C6%2C42%2C43%2C328%2C330%2C322%2C323%2C326%2C332%2C31%2C30%2C29&season=2021&month=1000&season1=2015&ind=0&team=&rost=&age=&filter=&players=&startdate={}&enddate={}&page=1_5000&pagenum=1&pageitems=2000000000".format(start_date, end_date) PITCHERS_URL = "https://www.fangraphs.com/leaders-legacy.aspx?pos=all&stats=pit&lg=all&qual=0&type=c,13,7,8,120,121,331,105,111,24,19,14,329,324,45,122,6,42,43,328,330,322,323,326,332,31,30,29&season=2021&month=1000&season1=2015&ind=0&team=&rost=&age=&filter=&players=&startdate={}&enddate={}&page=1_2000".format(start_date, end_date) # request the data pitchers_html = requests.get(PITCHERS_URL).text soup = BeautifulSoup(pitchers_html, "lxml") table = soup.find("table", {"class": "rgMasterTable"}) # get headers headers_html = table.find("thead").find_all("th") headers = [] for header in headers_html: headers.append(header.text) # get rows rows = [] rows_html = table.find("tbody").find_all("tr") for row in rows_html: row_data = [] for cell in row.find_all("td"): row_data.append(cell.text) rows.append(row_data) return pd.DataFrame(rows, columns = headers) def calc_speX(df, IP_limit): # A bunch of calculations return speX sdate = '2023-03-30' enddate = '2023-08-17' IP = 0 date_format = "%Y-%m-%d" start_date = datetime.datetime.strptime(sdate, date_format) end_date = datetime.datetime.strptime(enddate, date_format) daily = input("Do you want the daily values? ") writer = pd.ExcelWriter('speX-daily.xlsx', engine='openpyxl') if daily.lower()=="y": for single_date in pd.date_range(start=start_date, end=end_date): date_str = single_date.strftime(date_format) speX = parse_array_from_fangraphs_html(date_str, date_str) result = calc_speX(speX, IP) result.to_excel(writer, sheet_name=date_str) writer.save() writer.close()
但网站改版后的新版排行榜链接无法用requests直接爬取,尝试用Selenium渲染动态内容也失败,对应的Selenium代码如下:
chrome_options = Options() chrome_options.add_argument("--headless") chromedriver_path = '/chd/chromedriver.exe' driver = webdriver.Chrome(executable_path=chromedriver_path, options=chrome_options) driver.get(url) time.sleep(5) html = driver.page_source driver.quit() soup = BeautifulSoup(html, "lxml") table = soup.find("table", {"class": "rgMasterTable"})
运行后获取到的table为空。
可能的原因
- 新版页面的表格class名已变更,不再使用
rgMasterTable - 新版页面启用了反爬机制,检测到headless浏览器并限制内容加载
- 表格数据通过异步API加载,未直接渲染在初始HTML中
解决建议
1. 确认新版页面的表格选择器
打开新版页面的开发者工具(F12),在Elements面板中定位表格元素:
- 查找表格对应的真实class、id或其他属性,替换Soup中的查找条件
- 若页面用div模拟表格结构,需调整解析逻辑
2. 优化Selenium配置绕反爬
更新Selenium参数,模拟真实浏览器环境:
from selenium.webdriver.chrome.options import Options from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC chrome_options = Options() # 使用新版headless模式,更接近真实浏览器 chrome_options.add_argument("--headless=new") # 禁用自动化检测特征 chrome_options.add_argument("--disable-blink-features=AutomationControlled") # 设置真实用户代理 chrome_options.add_argument("--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/117.0.0.0 Safari/537.36") chrome_options.add_experimental_option("excludeSwitches", ["enable-automation"]) chrome_options.add_experimental_option('useAutomationExtension', False) chromedriver_path = '/chd/chromedriver.exe' driver = webdriver.Chrome(executable_path=chromedriver_path, options=chrome_options) driver.get(url) # 改用显式等待,直到表格加载完成(替换time.sleep) try: # 替换"new-table-class"为实际表格class wait = WebDriverWait(driver, 15) table_element = wait.until(EC.presence_of_element_located((By.CLASS_NAME, "new-table-class"))) html = driver.page_source finally: driver.quit() soup = BeautifulSoup(html, "lxml") table = soup.find("table", {"class": "new-table-class"})
3. 抓包获取API接口
通过开发者工具Network面板抓包:
- 刷新新版页面,筛选XHR/Fetch请求
- 找到返回表格数据的JSON接口
- 复制接口的请求URL、headers和参数,用
requests直接调用获取数据,效率远高于Selenium - 注意保留必要的请求头(如Referer、Cookie等)以通过验证
4. 临时继续使用旧版链接
若旧版兼容链接仍可用,可暂时继续使用,定期检查链接有效性即可
内容的提问来源于stack exchange,提问作者Carlos Marcano
相关产品推荐
相关产品推荐

