如何用Requests替代Selenium下载NHS网站CSV?排查HTTP 404错误
Fixing 404 Errors When Downloading NHS GP Patient CSV Files with Requests
Let's break down why you're hitting 404s and how to fix this with just requests—no need for Selenium or urllib.
Common Causes of Your 404 Error
- Relative URL Mistakes: The CSV links you're scraping might be relative paths (e.g.,
/data/file.csv) instead of full URLs. If you try to request these directly, you're hitting a non-existent address. - Missing Session Context: The NHS site likely requires maintaining a session (with cookies) to authorize downloads. If you're making isolated requests for CSV files without carrying over cookies from the main page, the server rejects your request.
- Ignoring Redirects: Some download links might redirect to the actual file, and while
requestsfollows redirects by default, not using a session can break this flow.
Step-by-Step Solution with Requests
Here's a complete, tested script that handles session persistence, full URL construction, and efficient downloads:
import requests from bs4 import BeautifulSoup # Initialize a session to preserve cookies and headers across requests session = requests.Session() session.headers.update({ 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' }) BASE_URL = "https://digital.nhs.uk" MAIN_PAGE = f"{BASE_URL}/data-and-information/publications/statistical/patients-registered-at-a-gp-practice" # 1. Scrape all monthly report links from the main page try: main_response = session.get(MAIN_PAGE) main_response.raise_for_status() # Fail fast if main page request fails except requests.exceptions.RequestException as e: print(f"Failed to access main page: {e}") exit() main_soup = BeautifulSoup(main_response.text, "html.parser") monthly_report_links = [] # Target the correct heading class (adjust if page structure changes) for heading in main_soup.find_all("h3", class_="nhsuk-heading-xs"): report_link = heading.find("a") if report_link and "Patients Registered at a GP Practice" in report_link.text: # Convert relative link to full URL full_link = BASE_URL + report_link["href"] monthly_report_links.append(full_link) # 2. Extract CSV download links from each report page csv_download_map = {} for report_url in monthly_report_links: try: report_response = session.get(report_url) report_response.raise_for_status() except requests.exceptions.RequestException as e: print(f"Failed to access report page {report_url}: {e}") continue report_soup = BeautifulSoup(report_response.text, "html.parser") csv_link = None # Find the CSV download link (look for "CSV" text or .csv in the URL) for link in report_soup.find_all("a", href=True): if link.text.strip() == "CSV" or link["href"].lower().endswith(".csv"): # Ensure we have a full URL if link["href"].startswith("http"): csv_link = link["href"] else: csv_link = BASE_URL + link["href"] break if csv_link: # Generate a clean filename from the report URL or title filename = report_url.split("/")[-1].replace("-", " ") + ".csv" csv_download_map[filename] = csv_link # 3. Download each CSV file efficiently for filename, csv_url in csv_download_map.items(): print(f"Starting download: {filename}") try: # Use stream=True to handle large files without loading them into memory download_response = session.get(csv_url, stream=True) download_response.raise_for_status() with open(filename, "wb") as f: for chunk in download_response.iter_content(chunk_size=8192): f.write(chunk) print(f"Successfully downloaded: {filename}") except requests.exceptions.RequestException as e: print(f"Failed to download {filename}: {e}")
Key Fixes Explained
- Session Persistence: The
requests.Session()object keeps track of cookies set by the NHS site when you first load the main page. This ensures your download requests are authorized. - Full URL Construction: We explicitly convert relative links to full URLs using the base NHS domain—this eliminates 404s from malformed requests.
- Streamed Downloads: Using
stream=Trueand writing chunks to the file makes the download efficient, even for large CSV files. - Error Handling: Added
raise_for_status()to catch HTTP errors early, making debugging easier.
Why You Don't Need urllib
requests is a more powerful, user-friendly alternative to urllib for this task. It handles sessions, headers, redirects, and streaming out of the box—no need to juggle separate libraries.
内容的提问来源于stack exchange,提问作者QHarr
相关产品推荐
相关产品推荐

