You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

抓取触发POST请求的网页PDF:带Token请求返回OK却无文件问题

How to Scrape PDFs Requiring Tokenized POST Requests

Got it, let's figure out how to grab those PDFs when the download button fires a tokenized POST request, and you're getting an OK response but no actual file. Here's a step-by-step breakdown of what you need to do:

1. Capture the Full Request Details First

First, fire up your browser's DevTools (F12) and head to the Network tab. Trigger the PDF download button, then find the POST request that gets sent. You need to note every detail of this request—missing even one small part will cause the server to return an empty/OK response instead of the PDF:

  • Request Headers: Pay close attention to Cookie, Referer, User-Agent, and Content-Type. Servers often use these to validate that the request is coming from a legitimate session/browser.
  • Form Data/Payload: Identify the token parameter (it might be named something like __RequestVerificationToken or token) and any other parameters tied to the specific PDF (like a file_id, decision_id, etc.). These parameters are critical—you can't just send the token alone.
  • Request URL: Make sure you have the exact endpoint URL the POST request is sent to (it's probably different from the list page you're viewing).

2. Extract the Token from the Initial Page

Tokens are almost always tied to your current session, and they're usually embedded in the HTML of the list page as a hidden input field. You need to fetch this page first, extract the token, and keep the session alive (so the server recognizes your request later).

Here's a quick Python example using requests and BeautifulSoup:

import requests
from bs4 import BeautifulSoup

# Use a Session to persist cookies across requests
session = requests.Session()

# The initial list page URL you provided
initial_page_url = "https://dsscic.nic.in/cause-list-report-web/view-decision?commissionname=302&file_category=1&fileno=&name=&public_authority=&decisiontypeid=1&frdate=&todate=&page_length=10&search_button=Submit"

# Fetch the page to get the token and session cookies
response = session.get(initial_page_url)
soup = BeautifulSoup(response.text, "html.parser")

# Replace with the actual name of the token input field (check DevTools!)
token = soup.find("input", {"name": "__RequestVerificationToken"})["value"]

3. Build the Exact POST Request

Now, construct the POST request using the same session (to keep cookies), the token you extracted, and all the required parameters from the DevTools capture. Don't skip any headers or payload fields—servers are picky about this.

Example continuation of the code above:

# Replace with the actual POST endpoint URL from DevTools
download_endpoint = "https://dsscic.nic.in/cause-list-report-web/your-download-endpoint"

# Payload must include ALL parameters from the captured POST request
payload = {
    "__RequestVerificationToken": token,
    "file_id": "12345",  # Replace with the specific PDF's ID from the list page/DevTools
    "decisiontypeid": "1",
    # Add any other parameters you saw in the Form Data section
}

# Headers should match what your browser sends
headers = {
    "Referer": initial_page_url,
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36",
    # Add any other headers like X-Requested-With if present in the captured request
}

# Send the POST request
pdf_response = session.post(download_endpoint, data=payload, headers=headers)

4. Handle the Response Properly

Once you get the response, check if it's actually a PDF:

# Verify the response is a PDF
if pdf_response.status_code == 200 and "application/pdf" in pdf_response.headers.get("Content-Type", ""):
    # Write the content to a file
    with open("downloaded_decision.pdf", "wb") as f:
        f.write(pdf_response.content)
    print("PDF downloaded successfully!")
else:
    # If not, print details to debug
    print(f"Status Code: {pdf_response.status_code}")
    print(f"Content-Type: {pdf_response.headers.get('Content-Type')}")
    print("Server Response Text:")
    print(pdf_response.text)

5. Troubleshooting Common Issues

If you're still getting an OK response but no PDF, check these:

  • Token validity: Tokens are often one-time use or session-bound—make sure you're using a fresh token from the same session that made the initial GET request.
  • Missing parameters: Double-check that you included every single parameter from the captured POST request (even ones that seem irrelevant).
  • Cookie persistence: Using a Session object is crucial here—don't make separate GET and POST requests without it, or the server won't recognize your session.
  • Anti-scraping measures: Some servers block non-browser User-Agents, or require you to simulate the full search flow (like submitting the search form before accessing the download links).

内容的提问来源于stack exchange,提问作者Nischay Namdev

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:35:46