You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python新手求助:如何爬取JS动态加载的英国赛马结果数据

Hey there! As someone who's tackled dynamic content scraping and worked with large datasets before, let's break down how you can reach your goal of scraping that horse racing data—especially since you're new to Python, we'll keep this practical and easy to follow.

1. First, Understand the Dynamic Loading

The site uses JS to load data, which means the content isn't in the initial HTML you get when you send a basic request. Instead, the browser sends additional requests (usually to an API) to fetch the data as JSON, then renders it on the page. Your job is to find those API requests instead of trying to scrape the rendered HTML directly.

Here's how to find them:

  • Open your browser's dev tools (press F12 or right-click > Inspect)
  • Go to the Network tab, then filter by "XHR" or "Fetch"
  • Refresh the results page—you'll see a list of requests. Look for ones that return JSON data with race results (check the "Response" tab to preview the content)
  • Note the request URL, parameters (like year, racecourse, page number), and headers (especially the User-Agent—this tells the site you're a real browser)
2. Tools You'll Need (All Beginner-Friendly)
  • requests: To send HTTP requests to the API and get the JSON data
  • pandas: To organize and store the large dataset (way easier than handling lists/dicts manually)
  • time: To add delays between requests (so you don't get blocked by the site)
  • Optional: selenium—only if you can't find the API (but API scraping is faster and simpler, so start here first)

Install them via pip if you haven't:

pip install requests pandas
3. Step-by-Step Implementation

Let's start with a small test before scaling to 20 years of data.

a. Test the API Request

First, copy the API URL and parameters you found in dev tools. For example, if the request looks like https://www.britishhorseracing.com/api/results?year=2024&racecourse=Ascot&page=1, you can test it with requests:

import requests

# Set headers to mimic a browser
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/124.0.0.0 Safari/537.36"
}

# Example API request (replace with the actual URL/params you found)
params = {
    "year": 2024,
    "racecourse": "Ascot",
    "page": 1
}

response = requests.get("YOUR_API_URL_HERE", headers=headers, params=params)
data = response.json()

# Print the first race result to check
print(data["results"][0])

b. Extract and Store Data

Once you confirm the data is coming through, extract the fields you care about (race date, horse name, position, time, etc.) and save it to a CSV:

import pandas as pd

# Extract relevant fields from the JSON
race_results = []
for race in data["results"]:
    for horse in race["horses"]:
        race_results.append({
            "date": race["date"],
            "racecourse": race["racecourse"],
            "horse_name": horse["name"],
            "position": horse["position"],
            "finish_time": horse["finish_time"]
        })

# Convert to DataFrame and save to CSV
df = pd.DataFrame(race_results)
df.to_csv("2024_ascot_results.csv", index=False)

c. Scale to Multiple Years/Racecourses

To handle 20 years of data, you'll need to loop through years, racecourses, and pages (since results are paginated). Add delays and error handling to avoid blocks:

import time

# List of years (adjust as needed)
years = range(2004, 2025)
# List of racecourses (you can get this from the site's filters or API)
racecourses = ["Ascot", "Cheltenham", "Aintree"]

all_results = []

for year in years:
    for course in racecourses:
        page = 1
        while True:
            try:
                params = {"year": year, "racecourse": course, "page": page}
                response = requests.get("YOUR_API_URL_HERE", headers=headers, params=params)
                response.raise_for_status()  # Raise error if request fails
                data = response.json()
                
                if not data["results"]:  # No more results on this page
                    break
                
                # Extract data as before
                for race in data["results"]:
                    for horse in race["horses"]:
                        all_results.append({
                            "year": year,
                            "racecourse": course,
                            "date": race["date"],
                            "horse_name": horse["name"],
                            "position": horse["position"],
                            "finish_time": horse["finish_time"]
                        })
                
                page += 1
                time.sleep(1)  # Wait 1 second between requests to avoid blocks
                
            except Exception as e:
                print(f"Error fetching {year} {course} page {page}: {e}")
                time.sleep(5)  # Wait longer before retrying

# Save all data to a single CSV
final_df = pd.DataFrame(all_results)
final_df.to_csv("all_horse_racing_results.csv", index=False)
4. Key Tips to Avoid Getting Blocked
  • Use a real User-Agent: Don't leave it blank—sites block requests without this.
  • Add delays: Even 1-2 seconds between requests goes a long way.
  • Start small: Test with 1 year and 1 racecourse first, then scale up.
  • Check the site's terms: Make sure scraping is allowed (most sites prohibit commercial use, but personal learning is usually okay).

Remember, as a beginner, it's okay to take it slow. If you hit a snag with finding the API, feel free to dig deeper into the dev tools—look for requests that update when you change filters (like year or racecourse) on the site.

内容的提问来源于stack exchange,提问作者M. Sprout

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 11:20:09