Web爬虫隐藏API:请求方式、XSRF-TOKEN及请求头选择问询
Answers to Your Web Scraping Questions
1. Is fetching the webpage first then requesting the JSON with cookies better (more like real user behavior)?
Absolutely—this approach is far more aligned with how a real user interacts with the site. Here’s why:
- When a user visits the page, their browser first loads the HTML, which sets necessary cookies (like
JSESSIONIDandXSRF-TOKEN) via the initial response. - The page’s JavaScript then fires the POST request to load the JSON, using those cookies and context-specific headers (like
Origin,Referer, and Fetch Metadata headers). - Directly hitting the JSON URL with a GET request skips this entire context setup. The request will have a different
Sec-Fetch-*profile (marked asnavigateinstead ofcors) and may lack cookies that are only set when loading the main page, making it easier for anti-scraping systems to flag as suspicious.
So yes, mimicking the full user flow (page load → AJAX POST) is the safer, more authentic choice.
2. Is the XSRF-TOKEN parameter required?
It depends on the request method and the backend’s validation rules, but here’s the typical behavior:
- For POST requests: Almost certainly required. XSRF tokens are designed to prevent cross-site request forgery, and most backends enforce this check for state-changing requests (even if this POST is just fetching data—many teams apply the rule universally). You’ll need to ensure the
X-XSRF-TOKENheader matches the value in theXSRF-TOKENcookie (both are set when loading the main page). - For GET requests: Maybe not, but it’s still wise to include it. Some backends skip XSRF validation for GET requests (since they’re supposed to be read-only), but including the token keeps your request identical to what a browser would send if someone manually navigated to the JSON URL.
3. Which request header fields should you prioritize including?
Focus on headers that signal authenticity and context—here’s a breakdown for each request type:
For the POST (AJAX) request:
- Mandatory fields:
Host: Target domain (critical for routing)User-Agent: Rotate this as planned—this is one of the easiest ways to avoid being flaggedAccept: Match the exact value from your capture (application/json, text/plain, */*)X-XSRF-TOKEN: Must match theXSRF-TOKENcookie valueContent-Type:application/json;charset=UTF-8(tells the backend you’re sending JSON)Origin: The domain of the main page (ensures same-origin context)Referer: The full URL of the main page that triggered the AJAX requestCookie: IncludeJSESSIONIDandXSRF-TOKEN(both from the initial page load)
- Highly recommended fields:
Connection: keep-alive(browser default behavior)Sec-Fetch-Site: same-origin,Sec-Fetch-Mode: cors,Sec-Fetch-Dest: empty: These Fetch Metadata headers are sent by modern browsers and help prove your request isn’t coming from a bot or cross-site contextAccept-Encoding: gzip, deflate, br: Reduces payload size and matches browser behaviorAccept-Language: en,en-US: Matches the user’s language setting from your capture
- Optional but authentic:
ADRUM: isAjax:true: This is a New Relic monitoring header—if the backend doesn’t block it, including it adds extra realism
For the GET request (if you ever use it):
- Mandatory fields:
HostUser-AgentCookie:JSESSIONIDandXSRF-TOKEN(same as above)
- Highly recommended fields:
Accept: Match the captured value (text/html,application/xhtml+xml,...)Sec-Fetch-Site: none,Sec-Fetch-Mode: navigate,Sec-Fetch-Dest: document: Reflects a direct navigation to the URLAccept-EncodingandAccept-Language(same as POST)
- Optional:
Connection: keep-aliveUpgrade-Insecure-Requests: 1: Only relevant if the site uses HTTPS, but it’s safe to include
内容的提问来源于stack exchange,提问作者Jo123
相关产品推荐
相关产品推荐

