You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python修改BARB网站过滤器,批量下载指定时段Top100节目数据?

Hey there! Based on your goal to scrape weekly Top100 data from the BARB site since Jan 2016, here's a step-by-step approach tailored to your experience with requests and Beautiful Soup. The key is to work with the site's underlying API (instead of parsing the HTML form directly) because that's more reliable for automated downloads.

Step 1: Inspect the Site's Request Flow

First, we need to figure out how the filters and download button work under the hood. Here's how to do it:

  • Open the target URL in Chrome/Firefox, then launch DevTools (F12 or Ctrl+Shift+I).
  • Go to the Network tab and check "Preserve log" to keep requests visible when the page refreshes.
  • Adjust the filters (Year=2016, Month=January, Week=1, Activity=Live streaming) and click "Apply" or the "Download CSV" button.
  • Look for the POST request that triggers the data load/download. Click on it to view:
    • The Request URL (this is our target API endpoint).
    • The Form Data (the exact parameters sent for filters: year, month, week, activity, etc.).
    • The Response—if it’s a CSV file, perfect! If it’s JSON, we can convert it to CSV later.
Step 2: Set Up Your Environment

Make sure you have the necessary libraries installed:

pip install requests python-dotenv beautifulsoup4

We’ll use requests.Session() to maintain a persistent session (critical for keeping cookies/session state) and avoid getting blocked.

Step 3: Write the Core Scraper Code

Here’s a sample script that replicates the CSV download request. You’ll need to replace placeholder values with what you found in DevTools:

import requests
import time
from datetime import datetime, timedelta
from bs4 import BeautifulSoup

# Initialize a session to persist cookies across requests
session = requests.Session()

# First, visit the main page to grab any required session cookies/tokens
main_url = "http://www.barb.co.uk/project-dovetail/top-100-programmes-broadcasters-own-player-apps/"
response = session.get(main_url)

# Extract CSRF token if required (check Form Data in DevTools for this)
csrf_token = None
soup = BeautifulSoup(response.text, "html.parser")
csrf_input = soup.find("input", {"name": "csrf_token"})
if csrf_input:
    csrf_token = csrf_input["value"]

# Replace with the CSV download endpoint you found in DevTools
csv_download_url = "http://www.barb.co.uk/path/to/csv/endpoint"

def download_weekly_data(year, month, week, activity):
    """Download CSV data for a specific week and activity type."""
    # Build the payload with exact parameters from DevTools
    payload = {
        "year": str(year),
        "month": str(month),
        "week": str(week),
        "activity": activity,  # e.g., "live_streaming" or "on_demand" (match DevTools values)
        "format": "csv",
        # Add other required parameters here (e.g., "filter_category", "csrf_token")
    }
    if csrf_token:
        payload["csrf_token"] = csrf_token
    
    try:
        response = session.post(csv_download_url, data=payload)
        response.raise_for_status()  # Trigger error for HTTP codes >=400
        
        # Create a descriptive filename
        filename = f"top100_{year}_{month:02d}_week{week}_{activity}.csv"
        with open(filename, "wb") as f:
            f.write(response.content)
        
        print(f"✅ Saved {filename}")
        return True
    except Exception as e:
        print(f"❌ Failed to download {year}-{month}-week{week} {activity}: {str(e)}")
        return False

def generate_weeks(start_date, end_date):
    """Generate year, month, week number for each week between two dates."""
    current_date = start_date
    while current_date <= end_date:
        # Get ISO week (adjust if site uses different week start, e.g., Sunday)
        iso_year, iso_week, _ = current_date.isocalendar()
        # Use the month of the week's start date to match site logic
        week_start = current_date - timedelta(days=current_date.weekday())  # Monday start
        month = week_start.month
        
        yield iso_year, month, iso_week
        current_date += timedelta(weeks=1)

# Main execution
start_date = datetime(2016, 1, 1)
end_date = datetime.now()

for year, month, week in generate_weeks(start_date, end_date):
    # Download both Live and On-demand data for each week
    download_weekly_data(year, month, week, "live_streaming")
    download_weekly_data(year, month, week, "on_demand")
    
    # Add delay to avoid rate limiting (increase to 5-10s if you get blocked)
    time.sleep(2)
Step 4: Key Troubleshooting Tips
  • Parameter Exactness: Double-check that the payload keys and values match exactly what you saw in DevTools. For example, if the site uses activity=1 instead of "live_streaming", update that.
  • Week Numbering: The script uses ISO weeks (Monday start). If the site uses Sunday-start weeks, adjust the week_start calculation to current_date - timedelta(days=current_date.weekday() + 1).
  • Rate Limiting: If you get blocked, increase the time.sleep() value or add random delays (e.g., time.sleep(random.uniform(2,5))).
  • Missing Data: Some weeks (like the current incomplete week) might not have data— the try-except block will log these errors without breaking the script.
Step 5: Validate Your Data

Spot-check a few downloaded CSVs against manual downloads from the site to ensure the data matches. This helps catch issues with parameter values or week numbering.

内容的提问来源于stack exchange,提问作者OD1995

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 04:57:15