无法抓取AJAX驱动网站的XHR/JSON数据怎么办?——Barchart页面数据抓取方案及同类网站排查要点
Hey there, let's work through this— I’ve tangled with Barchart’s AJAX-powered data loads more than a few times, so I know exactly where you’re stuck. Let’s start with fixing your current issue, then cover the key things to watch for with these kinds of sites going forward.
Fixing the Barchart ESU22 Options Data Grab
First off, the problem is likely that you’re targeting the wrong endpoint. The table on that page doesn’t load data directly with the page request—it pulls it from a dedicated API endpoint that you’ll need to track down in your browser’s dev tools. Here’s how to do it, plus a working code example:
Find the right API endpoint
Open Chrome/Firefox DevTools (F12), go to the Network tab, filter for XHR/Fetch, then refresh the page. Look for a request that returns JSON with the option strike prices, volumes, etc.—for your page, it’ll be something likehttps://www.barchart.com/proxies/core-api/v1/options/get.Copy the critical parameters and headers
Click into that request to see its details:- The
Paramstab will have required values likesymbol(ESU22),expirationDate(you’ll need to convert MI7N22 to the actual date—for this contract, it’s 2022-09-16), and fields specifying what data to pull. - The
Headerstab shows which ones are necessary—you don’t need all of them, just the key ones likeUser-Agent,Referer, andX-Requested-With.
- The
Working Python code to pull the data into a DataFrame
Here’s a snippet usingrequestsandpandasthat should work (tweak parameters if needed):import requests import pandas as pd # The actual API endpoint for options data api_url = "https://www.barchart.com/proxies/core-api/v1/options/get" # Parameters matching what the browser sends request_params = { "symbol": "ESU22", "expirationDate": "2022-09-16", "fields": "strike,lastPrice,change,changePercent,volume,openInterest,bid,ask", "groupBy": "strike", "orderBy": "strike", "orderDir": "asc", "meta": "field.shortName,field.type,field.description", "raw": "1" } # Minimal required headers to avoid being blocked request_headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36", "Referer": "https://www.barchart.com/futures/quotes/ESU22/options/MI7N22", "X-Requested-With": "XMLHttpRequest" } # Send the request and parse the JSON response = requests.get(api_url, params=request_params, headers=request_headers) response.raise_for_status() # Will throw an error if the request fails option_data = response.json() # Convert the data to a DataFrame df = pd.DataFrame(option_data["data"]) print(df.head())
Key Lessons for Future AJAX Data Scraping
When you run into sites like this again, keep these steps in mind to save time:
Always start with the Network tab
Don’t waste time parsing HTML—AJAX sites almost always load data via API calls. Filter for XHR/Fetch, refresh, and look for requests that return structured JSON matching the page’s table.Don’t guess parameters—steal them from the browser
Sites like Barchart rely on specific parameters (like expiration dates converted to YYYY-MM-DD) to return the right data. Copy these directly from the successful request in DevTools instead of trying to build them from scratch.Only use the necessary headers
You don’t need to copy every single header from the browser. Focus on:User-Agent: Mimics a real browser to avoid being flagged as a botReferer: Proves your request is coming from the legitimate pageX-Requested-With: Tells the server it’s an AJAX request (some sites block requests without this)Cookie: Only needed if the site requires authentication or a session (for public data like this, you usually don’t need it)
Watch for dynamic API logic
If the API endpoint or parameters are generated by JavaScript (e.g., encrypted keys), use the Sources tab in DevTools to search for keywords related to the request (like "options/get") and trace how the request is built.Avoid Selenium unless you have to
You’re smart to skip Selenium here—direct API calls are faster, more reliable, and way easier to turn into DataFrames. Only reach for Selenium if the API is heavily obfuscated or requires user interactions (like clicking a "Load More" button) that can’t be replicated with requests.
内容的提问来源于stack exchange,提问作者Nandril

