如何用Python修改BARB网站过滤器,批量下载指定时段Top100节目数据?
Hey there! Based on your goal to scrape weekly Top100 data from the BARB site since Jan 2016, here's a step-by-step approach tailored to your experience with requests and Beautiful Soup. The key is to work with the site's underlying API (instead of parsing the HTML form directly) because that's more reliable for automated downloads.
First, we need to figure out how the filters and download button work under the hood. Here's how to do it:
- Open the target URL in Chrome/Firefox, then launch DevTools (F12 or Ctrl+Shift+I).
- Go to the Network tab and check "Preserve log" to keep requests visible when the page refreshes.
- Adjust the filters (Year=2016, Month=January, Week=1, Activity=Live streaming) and click "Apply" or the "Download CSV" button.
- Look for the POST request that triggers the data load/download. Click on it to view:
- The Request URL (this is our target API endpoint).
- The Form Data (the exact parameters sent for filters: year, month, week, activity, etc.).
- The Response—if it’s a CSV file, perfect! If it’s JSON, we can convert it to CSV later.
Make sure you have the necessary libraries installed:
pip install requests python-dotenv beautifulsoup4
We’ll use requests.Session() to maintain a persistent session (critical for keeping cookies/session state) and avoid getting blocked.
Here’s a sample script that replicates the CSV download request. You’ll need to replace placeholder values with what you found in DevTools:
import requests import time from datetime import datetime, timedelta from bs4 import BeautifulSoup # Initialize a session to persist cookies across requests session = requests.Session() # First, visit the main page to grab any required session cookies/tokens main_url = "http://www.barb.co.uk/project-dovetail/top-100-programmes-broadcasters-own-player-apps/" response = session.get(main_url) # Extract CSRF token if required (check Form Data in DevTools for this) csrf_token = None soup = BeautifulSoup(response.text, "html.parser") csrf_input = soup.find("input", {"name": "csrf_token"}) if csrf_input: csrf_token = csrf_input["value"] # Replace with the CSV download endpoint you found in DevTools csv_download_url = "http://www.barb.co.uk/path/to/csv/endpoint" def download_weekly_data(year, month, week, activity): """Download CSV data for a specific week and activity type.""" # Build the payload with exact parameters from DevTools payload = { "year": str(year), "month": str(month), "week": str(week), "activity": activity, # e.g., "live_streaming" or "on_demand" (match DevTools values) "format": "csv", # Add other required parameters here (e.g., "filter_category", "csrf_token") } if csrf_token: payload["csrf_token"] = csrf_token try: response = session.post(csv_download_url, data=payload) response.raise_for_status() # Trigger error for HTTP codes >=400 # Create a descriptive filename filename = f"top100_{year}_{month:02d}_week{week}_{activity}.csv" with open(filename, "wb") as f: f.write(response.content) print(f"✅ Saved {filename}") return True except Exception as e: print(f"❌ Failed to download {year}-{month}-week{week} {activity}: {str(e)}") return False def generate_weeks(start_date, end_date): """Generate year, month, week number for each week between two dates.""" current_date = start_date while current_date <= end_date: # Get ISO week (adjust if site uses different week start, e.g., Sunday) iso_year, iso_week, _ = current_date.isocalendar() # Use the month of the week's start date to match site logic week_start = current_date - timedelta(days=current_date.weekday()) # Monday start month = week_start.month yield iso_year, month, iso_week current_date += timedelta(weeks=1) # Main execution start_date = datetime(2016, 1, 1) end_date = datetime.now() for year, month, week in generate_weeks(start_date, end_date): # Download both Live and On-demand data for each week download_weekly_data(year, month, week, "live_streaming") download_weekly_data(year, month, week, "on_demand") # Add delay to avoid rate limiting (increase to 5-10s if you get blocked) time.sleep(2)
- Parameter Exactness: Double-check that the
payloadkeys and values match exactly what you saw in DevTools. For example, if the site usesactivity=1instead of "live_streaming", update that. - Week Numbering: The script uses ISO weeks (Monday start). If the site uses Sunday-start weeks, adjust the
week_startcalculation tocurrent_date - timedelta(days=current_date.weekday() + 1). - Rate Limiting: If you get blocked, increase the
time.sleep()value or add random delays (e.g.,time.sleep(random.uniform(2,5))). - Missing Data: Some weeks (like the current incomplete week) might not have data— the try-except block will log these errors without breaking the script.
Spot-check a few downloaded CSVs against manual downloads from the site to ensure the data matches. This helps catch issues with parameter values or week numbering.
内容的提问来源于stack exchange,提问作者OD1995

