You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何快速下载在线表格数据?含芝加哥马拉松历史数据批量下载需求

How to Batch Download 1996-2016 Chicago Marathon Results (400k+ Entries)

Great question—dealing with paginated large datasets like this is a common pain point. Here are a few practical approaches, ordered from easiest to more technical:

1. First: Check for an Official Bulk Export Option

Before jumping into scraping, take 5 minutes to poke around the site:

  • Look for links labeled "Bulk Download", "Export All", or "Data Archive" in the footer, race results page menus, or under a "Resources" section.
  • Some timing platforms (like Mikatiming, which powers this site) often hide bulk data options for registered users or researchers—worth checking if a login unlocks this feature.

If you find this, it’s by far the best method (no scraping needed, and you’ll get clean, structured data).

2. No-Code Browser Extension (Quickest for Non-Developers)

If there’s no official export, browser extensions built for scraping paginated tables work perfectly here:

  • Data Miner or Scraper (both available for Chrome/Firefox):
    1. Install the extension and navigate to the first results page of your target year.
    2. Select a single runner’s table row—most extensions will auto-detect the repeating pattern for all entries.
    3. Configure pagination: tell the extension how to navigate to the next page (either by clicking the "Next" button, or incrementing the page number in the URL).
    4. Set the export format to CSV or Excel, then let the extension run. It will loop through all pages, collect data, and save it as a single file.
  • Pro tip: Add a 1-2 second delay between page loads to avoid triggering anti-scraping measures.

3. Python Script (Most Flexible for Developers)

For full control over the process, write a simple Python script to automate scraping. Here’s a basic example using requests and pandas (install them first with pip install requests pandas):

import requests
import pandas as pd
import time

# Base URL pattern—adjust based on how the site paginates (check URL when clicking "Next")
base_url = "http://chicago-history.r.mikatiming.de/{year}/?page={page}"
years = range(1996, 2017)
all_results = pd.DataFrame()

for year in years:
    print(f"Starting year {year}...")
    # Calculate total pages (400k total / 1000 per page = ~400 pages per year)
    for page_num in range(1, 401):
        try:
            response = requests.get(
                base_url.format(year=year, page=page_num),
                headers={"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36"}
            )
            response.raise_for_status()
            
            # Parse HTML table into DataFrame—adjust index if multiple tables exist
            page_data = pd.read_html(response.text)[0]
            all_results = pd.concat([all_results, page_data], ignore_index=True)
            
            time.sleep(1)  # Polite delay to avoid server blocks
            
        except Exception as e:
            print(f"Failed to scrape page {page_num} of {year}: {str(e)}")
            continue

# Save all data to a single CSV
all_results.to_csv("chicago_marathon_1996-2016_full.csv", index=False)
print("Scraping complete! Data saved to chicago_marathon_1996-2016_full.csv")
  • Notes:
    • Verify the pagination parameter (some sites use offset instead of page—check the URL when navigating between pages).
    • If the site blocks requests, add cookies from your browser session to the requests.get() call.
    • Always check the site’s robots.txt file to confirm scraping is allowed.

4. Command-Line Tools (For CLI Enthusiasts)

If you prefer the command line, use wget to download all pages first, then parse the HTML to extract data:

  1. Download all pages for a year with a loop:
    for page in {1..400}; do wget -q "http://chicago-history.r.mikatiming.de/1996/?page=$page" -O 1996_page_$page.html; sleep 1; done
    
  2. Use tools like pup (HTML parser) or grep to extract table rows from all HTML files, then combine them into a CSV. This is more manual but works well if you’re comfortable with CLI workflows.

Key Considerations

  • Rate Limiting: Never hammer the server—1-2 second delays between requests are standard practice to avoid being blocked.
  • Terms of Service: Ensure scraping doesn’t violate the site’s terms of use. Most allow non-commercial, personal-use scraping, but it’s worth double-checking.
  • Data Cleaning: After scraping, you’ll likely need to clean the data (remove duplicates, fix formatting, etc.) using Excel or pandas.

内容的提问来源于stack exchange,提问作者hkmat999

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 04:22:23