You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Python爬取Racenet赛马历史数据的技术求助(日历交互问题)

解决方案:遍历日期获取Racenet赛马赛事URL

核心思路

这个网站的日期选择是动态AJAX加载,页面源码里看不到赛事链接,因为数据是通过API异步获取的。解决步骤如下:

  1. 抓包获取数据接口:用浏览器开发者工具找到选择日期后加载赛事列表的API接口
  2. 模拟API请求:用Python发送请求获取当日所有赛事的结构化数据
  3. 拼接赛事URL:从API返回的JSON数据中提取赛道标识、赛事标识等字段,拼接成完整的赛事URL
  4. 遍历日期范围:生成目标日期区间,逐个请求API收集所有URL

具体实现代码

1. 依赖库安装

先安装需要的库:

pip install requests python-dateutil

2. 完整代码

import requests
import datetime
import time

def get_daily_race_urls(target_date):
    # 替换为你抓包得到的实际API接口
    api_url = "https://www.racenet.com.au/api/results/meetings"
    # 模拟浏览器请求头,避免被拦截
    headers = {
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36",
        "Referer": "https://www.racenet.com.au/results/horse-racing",
        "Accept": "application/json, text/plain, */*"
    }
    # 日期参数格式要和API要求一致,一般是YYYY-MM-DD
    date_str = target_date.strftime("%Y-%m-%d")
    params = {
        "date": date_str
        # 如果抓包发现有其他参数(比如region),这里也要加上
    }

    try:
        response = requests.get(api_url, headers=headers, params=params)
        response.raise_for_status()  # 抛出HTTP错误
        meeting_data = response.json()
    except Exception as e:
        print(f"请求日期 {date_str} 失败: {str(e)}")
        return []

    race_urls = []
    # 遍历每个赛道的赛事
    for meeting in meeting_data.get("meetings", []):
        track_slug = meeting.get("slug")
        # 把日期转成YYYYMMDD格式,用于URL拼接
        date_formatted = date_str.replace("-", "")
        # 遍历该赛道的所有场次
        for race in meeting.get("races", []):
            race_slug = race.get("slug")
            race_number = race.get("number")
            # 拼接完整的赛事URL
            full_url = f"https://www.racenet.com.au/results/horse-racing/{track_slug}-{date_formatted}/{race_slug}-race-{race_number}"
            race_urls.append(full_url)
    return race_urls

# 定义要遍历的日期范围
start_date = datetime.date(2021, 5, 1)
end_date = datetime.date(2021, 5, 11)
date_delta = datetime.timedelta(days=1)

all_race_urls = []
current_date = start_date

# 遍历日期并收集URL
while current_date <= end_date:
    print(f"正在处理日期: {current_date.strftime('%Y-%m-%d')}")
    daily_urls = get_daily_race_urls(current_date)
    all_race_urls.extend(daily_urls)
    # 加延迟,避免触发反爬
    time.sleep(1.5)
    current_date += date_delta

# 保存结果到文件
with open("racenet_race_urls.txt", "w", encoding="utf-8") as f:
    for url in all_race_urls:
        f.write(f"{url}\n")

print(f"任务完成,共收集到 {len(all_race_urls)} 个赛事URL")

关键注意事项

  • 抓包确认API:网站的API可能会更新,一定要自己用浏览器开发者工具(F12 -> Network标签)抓包,查看选择日期后发送的XHR请求,确认API地址、参数和返回结构
  • 请求头模拟:必须带上User-Agent等关键请求头,否则服务器可能返回403或空数据;如果需要登录才能访问,要把浏览器的Cookie复制到请求头中
  • 反爬规避:不要频繁请求,加1-2秒的延迟;如果请求量很大,可以考虑使用代理IP
  • 异常处理:代码中加入了基础的异常捕获,你可以根据实际情况扩展,比如重试机制、日志记录等

内容的提问来源于stack exchange,提问作者JL89

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 17:43:24