You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python网页爬取问题:无法获取指定日期范围的MLB排名数据

解决MLB预测排名历史日期数据爬取问题

问题根源

你的代码无法获取指定日期范围数据的核心原因是URL参数格式错误:网站的日期筛选参数不是直接拼接日期字符串,而是需要通过date键值对传递。另外,直接用requests.get请求可能被网站识别为爬虫,返回默认日期页面。

修正方案

  1. 修正URL参数格式:将日期作为date参数传递,而非直接拼接在URL后
  2. 添加请求头模拟浏览器:避免被网站拦截或返回默认内容
  3. 优化日期循环逻辑:跳过无数据的日期(如休赛日),提升效率

修改后的完整代码

import requests
from bs4 import BeautifulSoup
import csv
from datetime import date, timedelta

# 目标URL
base_url = "https://www.teamrankings.com/mlb/ranking/predictive-by-other"

# 日期范围
start_date = date(2023, 3, 30)
end_date = date(2023, 10, 5)

csv_filename = "team_rankings.csv"
header = ["Date", "Rank", "Team", "Predictive Rank", "Rating"]

# 模拟浏览器请求头
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
}

with open(csv_filename, mode='w', newline='', encoding='utf-8') as file:
    writer = csv.writer(file)
    writer.writerow(header)

    current_date = start_date
    while current_date <= end_date:
        date_str = current_date.strftime('%Y-%m-%d')
        # 修正URL参数格式:使用date键传递日期
        current_url = f"{base_url}?date={date_str}"
        
        print(f"正在爬取: {date_str}")
        
        # 添加请求头发送请求
        response = requests.get(current_url, headers=headers)
        # 检查请求是否成功
        if response.status_code != 200:
            print(f"{date_str} 请求失败,状态码: {response.status_code}")
            current_date += timedelta(days=1)
            continue
            
        soup = BeautifulSoup(response.text, 'html.parser')
        table = soup.find('table', {'class': 'tr-table datatable scrollable'})

        if not table:
            print(f"{date_str} 无可用数据")
            current_date += timedelta(days=1)
            continue
            
        # 提取表格数据
        for row in table.find_all('tr')[1:]:
            cols = row.find_all('td')
            # 确保列数足够,避免索引错误
            if len(cols) >=4:
                rank = cols[0].text.strip()
                team = cols[1].text.strip()
                predictive_rank = cols[2].text.strip()
                rating = cols[3].text.strip()
                writer.writerow([date_str, rank, team, predictive_rank, rating])
        
        current_date += timedelta(days=1)

print("数据爬取并保存完成!")

关键修改点说明

  • URL参数:将原代码的?YYYY-MM-DD改为?date=YYYY-MM-DD,符合网站的参数规则
  • 请求头:添加User-Agent模拟浏览器,避免被网站限制
  • 错误处理:增加状态码检查和空表格判断,跳过无效日期,避免程序崩溃
  • 编码设置:打开CSV时指定encoding='utf-8',防止特殊字符乱码

内容的提问来源于stack exchange,提问作者Mike Friedrich

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.29 04:54:54