You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python自动获取JS加载页面的XHR请求URL及识别目标请求

Nice questions—scraping SEC EDGAR’s dynamic pages can feel like hunting for a needle in a haystack, so let’s walk through solutions for both of your concerns.

1. Automatically fetch the target XHR URL from the initial page URL

The initial URL you’re using is SEC’s IX viewer page, which loads the actual filing via an XHR call. The good news is you don’t need to load the whole page to get the target URL—you can parse the query parameter directly, or use a headless browser if you want a more robust fallback.

Method 1: Parse the query parameter (most efficient)

The initial URL’s doc query parameter directly points to the full path of the filing. You can extract this and build the target URL in just a few lines:

from urllib.parse import urlparse, parse_qs

initial_url = "https://www.sec.gov/ix?doc=/Archives/edgar/data/320193/000032019319000076/a10-qq320196292019.htm"
parsed_url = urlparse(initial_url)
# Extract the "doc" parameter value
doc_path = parse_qs(parsed_url.query)["doc"][0]
# Build the full target URL
target_url = f"https://www.sec.gov{doc_path}"

print(target_url)
# Output: https://www.sec.gov/Archives/edgar/data/320193/000032019319000076/a10-qq320196292019.htm

This works because SEC’s IX viewer is just a wrapper—all the actual filing content lives at the path specified in the doc parameter. No need to load JavaScript or wait for XHR calls here.

Method 2: Use a headless browser (for dynamic edge cases)

If SEC ever changes how their viewer works, you can use a headless browser like Playwright to load the page and capture XHR requests as they fire:

from playwright.sync_api import sync_playwright

initial_url = "https://www.sec.gov/ix?doc=/Archives/edgar/data/320193/000032019319000076/a10-qq320196292019.htm"
target_xhr_url = None

with sync_playwright() as p:
    # Launch headless Chrome
    browser = p.chromium.launch(headless=True)
    page = browser.new_page()
    
    # Listen for all network requests and filter XHRs
    def capture_target_request(request):
        nonlocal target_xhr_url
        # Check if it's an XHR and matches the filing path pattern
        if request.resource_type == "xhr" and "/Archives/edgar/data/" in request.url:
            target_xhr_url = request.url
    
    page.on("request", capture_target_request)
    # Wait until network is idle to ensure all XHRs fire
    page.goto(initial_url, wait_until="networkidle")
    browser.close()

print(target_xhr_url)

This method is slower but more flexible if the page’s dynamic behavior changes.

2. Distinguish the target XHR request from others

When dealing with multiple XHR calls, you can filter based on URL patterns, request/response metadata, and content. Here’s how to narrow it down:

Key filters to use:

  • URL signature: Target filings live under /Archives/edgar/data/ and include the company’s CIK (320193 for Apple) and a filing-specific filename (like a10-qq320196292019.htm for a 10-Q).
  • Request method: The target XHR will almost always be a GET request.
  • Response content type: The filing returns HTML, so check for text/html or application/xhtml+xml in the Content-Type header.
  • Response content: You can peek at the response body to confirm it contains filing-specific text (e.g., "Quarterly Report" or "10-Q").

Example filter logic in code

If using Playwright, you can refine the request listener to only pick the right XHR:

def filter_target_request(request):
    # Skip non-XHR requests
    if request.resource_type != "xhr":
        return False
    # Match the CIK and filing file extension
    if "/Archives/edgar/data/320193/" in request.url and request.url.endswith(".htm"):
        # Optional: Verify response content type after the request completes
        if request.response():
            content_type = request.response().headers.get("Content-Type", "")
            return "text/html" in content_type
    return False

You can also add a check for the response body if you want extra certainty:

if request.response():
    response_text = request.response().text()
    return "10-Q" in response_text or "Quarterly Report" in response_text

Important note: SEC anti-scraping rules

Make sure to set a valid User-Agent header (use your email address as part of it, per SEC’s guidelines) and avoid making too many requests too quickly. This helps keep your scraper compliant and avoids getting blocked.

内容的提问来源于stack exchange,提问作者moron

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 09:08:43