如何自动化从blob:https开头的URL下载美国CDC2022年猴痘疫情数据CSV文件
Hey there! Let's tackle this blob URL issue you're facing with the CDC monkeypox data. The core problem here is that blob URLs are browser-generated pointers to in-memory content—they aren't actual HTTP endpoints, which is why requests throws that InvalidSchema error, and direct Selenium GETs fail too.
Here are two reliable approaches to automate this data download without manual clicks:
1. Capture the Real Data API with Selenium DevTools
The blob link you see is created by frontend JavaScript after fetching data from a hidden API endpoint. We can use Selenium's DevTools to intercept network requests and grab that actual API URL.
Here's a working script:
from selenium import webdriver from selenium.webdriver.chrome.service import Service from webdriver_manager.chrome import ChromeDriverManager import requests import time import json # Set up Chrome with DevTools enabled options = webdriver.ChromeOptions() driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()), options=options) driver.execute_cdp_cmd('Network.enable', {}) # Track the data request we care about target_request = None def log_requests(request): global target_request # Filter requests that likely contain the monkeypox data url = request['url'] if ('monkeypox' in url.lower() or 'response' in url.lower()) and ('json' in url.lower() or 'csv' in url.lower()): target_request = request # Attach the request listener driver.execute_cdp_cmd('Network.requestWillBeSent', {'listener': log_requests}) # Load the CDC page driver.get('https://www.cdc.gov/poxvirus/monkeypox/response/2022/us-map.html') # Wait for the page to load and data to be fetched time.sleep(5) # Replace with WebDriverWait for better reliability if target_request: # Use requests to fetch the data directly from the real API headers = { 'User-Agent': driver.execute_script('return navigator.userAgent;'), 'Referer': 'https://www.cdc.gov/' } response = requests.get(target_request['url'], headers=headers) # Save the data (adjust format based on what the API returns) if target_request['url'].endswith('.json'): with open('monkeypox_data.json', 'w') as f: json.dump(response.json(), f, indent=2) else: # Assume CSV with open('monkeypox_data.csv', 'w') as f: f.write(response.text) print("Data downloaded successfully!") else: print("Couldn't find the data request—tweak the filter conditions in log_requests()") driver.quit()
2. Extract Data Directly from Page JavaScript
Sometimes the CDC embeds the data directly in a global JavaScript variable on the page. You can use Selenium to execute JS and pull that data out:
from selenium import webdriver from selenium.webdriver.chrome.service import Service from webdriver_manager.chrome import ChromeDriverManager import json import time driver = webdriver.Chrome(service=Service(ChromeDriverManager().install())) driver.get('https://www.cdc.gov/poxvirus/monkeypox/response/2022/us-map.html') # Wait for data to load into the page's JS context time.sleep(5) # Try to extract the data (you'll need to inspect the page's JS to find the right variable) # Example: If the data is stored in window.monkeypoxResponseData data = driver.execute_script('return window.monkeypoxResponseData || null;') if data: with open('monkeypox_data.json', 'w') as f: json.dump(data, f, indent=2) print("Data extracted from page JS successfully!") else: print("No matching global JS variable found—check the page's source code for the data storage") driver.quit()
Why Your Original Methods Failed
Blob URLs like blob:https://www.cdc.gov/... are only valid within the browser session that created them. They point to content stored in the browser's memory, not a public web server. That's why requests can't resolve them (it doesn't understand the blob: schema) and Selenium can't fetch them directly—there's no actual HTTP resource at that "URL".
Pro Tip
Once you capture the real API URL with the first method, you can often reuse it directly in requests without launching Selenium every time (as long as the API doesn't change frequently). Just make sure to include the same headers (User-Agent, Referer) that the browser sends to avoid being blocked.
内容的提问来源于stack exchange,提问作者Jigsaw

