如何用Python自动获取JS加载页面的XHR请求URL及识别目标请求
Nice questions—scraping SEC EDGAR’s dynamic pages can feel like hunting for a needle in a haystack, so let’s walk through solutions for both of your concerns.
1. Automatically fetch the target XHR URL from the initial page URL
The initial URL you’re using is SEC’s IX viewer page, which loads the actual filing via an XHR call. The good news is you don’t need to load the whole page to get the target URL—you can parse the query parameter directly, or use a headless browser if you want a more robust fallback.
Method 1: Parse the query parameter (most efficient)
The initial URL’s doc query parameter directly points to the full path of the filing. You can extract this and build the target URL in just a few lines:
from urllib.parse import urlparse, parse_qs initial_url = "https://www.sec.gov/ix?doc=/Archives/edgar/data/320193/000032019319000076/a10-qq320196292019.htm" parsed_url = urlparse(initial_url) # Extract the "doc" parameter value doc_path = parse_qs(parsed_url.query)["doc"][0] # Build the full target URL target_url = f"https://www.sec.gov{doc_path}" print(target_url) # Output: https://www.sec.gov/Archives/edgar/data/320193/000032019319000076/a10-qq320196292019.htm
This works because SEC’s IX viewer is just a wrapper—all the actual filing content lives at the path specified in the doc parameter. No need to load JavaScript or wait for XHR calls here.
Method 2: Use a headless browser (for dynamic edge cases)
If SEC ever changes how their viewer works, you can use a headless browser like Playwright to load the page and capture XHR requests as they fire:
from playwright.sync_api import sync_playwright initial_url = "https://www.sec.gov/ix?doc=/Archives/edgar/data/320193/000032019319000076/a10-qq320196292019.htm" target_xhr_url = None with sync_playwright() as p: # Launch headless Chrome browser = p.chromium.launch(headless=True) page = browser.new_page() # Listen for all network requests and filter XHRs def capture_target_request(request): nonlocal target_xhr_url # Check if it's an XHR and matches the filing path pattern if request.resource_type == "xhr" and "/Archives/edgar/data/" in request.url: target_xhr_url = request.url page.on("request", capture_target_request) # Wait until network is idle to ensure all XHRs fire page.goto(initial_url, wait_until="networkidle") browser.close() print(target_xhr_url)
This method is slower but more flexible if the page’s dynamic behavior changes.
2. Distinguish the target XHR request from others
When dealing with multiple XHR calls, you can filter based on URL patterns, request/response metadata, and content. Here’s how to narrow it down:
Key filters to use:
- URL signature: Target filings live under
/Archives/edgar/data/and include the company’s CIK (320193 for Apple) and a filing-specific filename (likea10-qq320196292019.htmfor a 10-Q). - Request method: The target XHR will almost always be a
GETrequest. - Response content type: The filing returns HTML, so check for
text/htmlorapplication/xhtml+xmlin theContent-Typeheader. - Response content: You can peek at the response body to confirm it contains filing-specific text (e.g., "Quarterly Report" or "10-Q").
Example filter logic in code
If using Playwright, you can refine the request listener to only pick the right XHR:
def filter_target_request(request): # Skip non-XHR requests if request.resource_type != "xhr": return False # Match the CIK and filing file extension if "/Archives/edgar/data/320193/" in request.url and request.url.endswith(".htm"): # Optional: Verify response content type after the request completes if request.response(): content_type = request.response().headers.get("Content-Type", "") return "text/html" in content_type return False
You can also add a check for the response body if you want extra certainty:
if request.response(): response_text = request.response().text() return "10-Q" in response_text or "Quarterly Report" in response_text
Important note: SEC anti-scraping rules
Make sure to set a valid User-Agent header (use your email address as part of it, per SEC’s guidelines) and avoid making too many requests too quickly. This helps keep your scraper compliant and avoids getting blocked.
内容的提问来源于stack exchange,提问作者moron

