抓取触发POST请求的网页PDF:带Token请求返回OK却无文件问题
Got it, let's figure out how to grab those PDFs when the download button fires a tokenized POST request, and you're getting an OK response but no actual file. Here's a step-by-step breakdown of what you need to do:
1. Capture the Full Request Details First
First, fire up your browser's DevTools (F12) and head to the Network tab. Trigger the PDF download button, then find the POST request that gets sent. You need to note every detail of this request—missing even one small part will cause the server to return an empty/OK response instead of the PDF:
- Request Headers: Pay close attention to
Cookie,Referer,User-Agent, andContent-Type. Servers often use these to validate that the request is coming from a legitimate session/browser. - Form Data/Payload: Identify the token parameter (it might be named something like
__RequestVerificationTokenortoken) and any other parameters tied to the specific PDF (like afile_id,decision_id, etc.). These parameters are critical—you can't just send the token alone. - Request URL: Make sure you have the exact endpoint URL the POST request is sent to (it's probably different from the list page you're viewing).
2. Extract the Token from the Initial Page
Tokens are almost always tied to your current session, and they're usually embedded in the HTML of the list page as a hidden input field. You need to fetch this page first, extract the token, and keep the session alive (so the server recognizes your request later).
Here's a quick Python example using requests and BeautifulSoup:
import requests from bs4 import BeautifulSoup # Use a Session to persist cookies across requests session = requests.Session() # The initial list page URL you provided initial_page_url = "https://dsscic.nic.in/cause-list-report-web/view-decision?commissionname=302&file_category=1&fileno=&name=&public_authority=&decisiontypeid=1&frdate=&todate=&page_length=10&search_button=Submit" # Fetch the page to get the token and session cookies response = session.get(initial_page_url) soup = BeautifulSoup(response.text, "html.parser") # Replace with the actual name of the token input field (check DevTools!) token = soup.find("input", {"name": "__RequestVerificationToken"})["value"]
3. Build the Exact POST Request
Now, construct the POST request using the same session (to keep cookies), the token you extracted, and all the required parameters from the DevTools capture. Don't skip any headers or payload fields—servers are picky about this.
Example continuation of the code above:
# Replace with the actual POST endpoint URL from DevTools download_endpoint = "https://dsscic.nic.in/cause-list-report-web/your-download-endpoint" # Payload must include ALL parameters from the captured POST request payload = { "__RequestVerificationToken": token, "file_id": "12345", # Replace with the specific PDF's ID from the list page/DevTools "decisiontypeid": "1", # Add any other parameters you saw in the Form Data section } # Headers should match what your browser sends headers = { "Referer": initial_page_url, "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36", # Add any other headers like X-Requested-With if present in the captured request } # Send the POST request pdf_response = session.post(download_endpoint, data=payload, headers=headers)
4. Handle the Response Properly
Once you get the response, check if it's actually a PDF:
# Verify the response is a PDF if pdf_response.status_code == 200 and "application/pdf" in pdf_response.headers.get("Content-Type", ""): # Write the content to a file with open("downloaded_decision.pdf", "wb") as f: f.write(pdf_response.content) print("PDF downloaded successfully!") else: # If not, print details to debug print(f"Status Code: {pdf_response.status_code}") print(f"Content-Type: {pdf_response.headers.get('Content-Type')}") print("Server Response Text:") print(pdf_response.text)
5. Troubleshooting Common Issues
If you're still getting an OK response but no PDF, check these:
- Token validity: Tokens are often one-time use or session-bound—make sure you're using a fresh token from the same session that made the initial GET request.
- Missing parameters: Double-check that you included every single parameter from the captured POST request (even ones that seem irrelevant).
- Cookie persistence: Using a
Sessionobject is crucial here—don't make separate GET and POST requests without it, or the server won't recognize your session. - Anti-scraping measures: Some servers block non-browser User-Agents, or require you to simulate the full search flow (like submitting the search form before accessing the download links).
内容的提问来源于stack exchange,提问作者Nischay Namdev

