You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何自动化从blob:https开头的URL下载美国CDC2022年猴痘疫情数据CSV文件

Hey there! Let's tackle this blob URL issue you're facing with the CDC monkeypox data. The core problem here is that blob URLs are browser-generated pointers to in-memory content—they aren't actual HTTP endpoints, which is why requests throws that InvalidSchema error, and direct Selenium GETs fail too.

Here are two reliable approaches to automate this data download without manual clicks:

1. Capture the Real Data API with Selenium DevTools

The blob link you see is created by frontend JavaScript after fetching data from a hidden API endpoint. We can use Selenium's DevTools to intercept network requests and grab that actual API URL.

Here's a working script:

from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from webdriver_manager.chrome import ChromeDriverManager
import requests
import time
import json

# Set up Chrome with DevTools enabled
options = webdriver.ChromeOptions()
driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()), options=options)
driver.execute_cdp_cmd('Network.enable', {})

# Track the data request we care about
target_request = None

def log_requests(request):
    global target_request
    # Filter requests that likely contain the monkeypox data
    url = request['url']
    if ('monkeypox' in url.lower() or 'response' in url.lower()) and ('json' in url.lower() or 'csv' in url.lower()):
        target_request = request

# Attach the request listener
driver.execute_cdp_cmd('Network.requestWillBeSent', {'listener': log_requests})

# Load the CDC page
driver.get('https://www.cdc.gov/poxvirus/monkeypox/response/2022/us-map.html')

# Wait for the page to load and data to be fetched
time.sleep(5)  # Replace with WebDriverWait for better reliability

if target_request:
    # Use requests to fetch the data directly from the real API
    headers = {
        'User-Agent': driver.execute_script('return navigator.userAgent;'),
        'Referer': 'https://www.cdc.gov/'
    }
    response = requests.get(target_request['url'], headers=headers)
    
    # Save the data (adjust format based on what the API returns)
    if target_request['url'].endswith('.json'):
        with open('monkeypox_data.json', 'w') as f:
            json.dump(response.json(), f, indent=2)
    else:  # Assume CSV
        with open('monkeypox_data.csv', 'w') as f:
            f.write(response.text)
    print("Data downloaded successfully!")
else:
    print("Couldn't find the data request—tweak the filter conditions in log_requests()")

driver.quit()

2. Extract Data Directly from Page JavaScript

Sometimes the CDC embeds the data directly in a global JavaScript variable on the page. You can use Selenium to execute JS and pull that data out:

from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from webdriver_manager.chrome import ChromeDriverManager
import json
import time

driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()))
driver.get('https://www.cdc.gov/poxvirus/monkeypox/response/2022/us-map.html')

# Wait for data to load into the page's JS context
time.sleep(5)

# Try to extract the data (you'll need to inspect the page's JS to find the right variable)
# Example: If the data is stored in window.monkeypoxResponseData
data = driver.execute_script('return window.monkeypoxResponseData || null;')

if data:
    with open('monkeypox_data.json', 'w') as f:
        json.dump(data, f, indent=2)
    print("Data extracted from page JS successfully!")
else:
    print("No matching global JS variable found—check the page's source code for the data storage")

driver.quit()

Why Your Original Methods Failed

Blob URLs like blob:https://www.cdc.gov/... are only valid within the browser session that created them. They point to content stored in the browser's memory, not a public web server. That's why requests can't resolve them (it doesn't understand the blob: schema) and Selenium can't fetch them directly—there's no actual HTTP resource at that "URL".

Pro Tip

Once you capture the real API URL with the first method, you can often reuse it directly in requests without launching Selenium every time (as long as the API doesn't change frequently). Just make sure to include the same headers (User-Agent, Referer) that the browser sends to avoid being blocked.

内容的提问来源于stack exchange,提问作者Jigsaw

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.27 19:32:47