如何抓取非JSON格式表单数据的Ajax请求内容?
Hey there! I totally get the frustration when Ajax requests don’t behave like the JSON examples you find in tutorials—this form data format is actually super common, just not the one everyone starts with. Let’s walk through exactly how to make this work.
First, Understand the Data Format
The form data you’re seeing (filters[targetValueMin]: 0, etc.) is URL-encoded form data (the application/x-www-form-urlencoded content type), not JSON. That means you don’t use the json parameter in requests—you use the data parameter instead, passing a dictionary that matches the exact key-value pairs from the request.
Step-by-Step Implementation
Here’s a concrete example using Python’s requests library (the go-to for web scraping):
1. Grab the Correct Ajax Request URL
First, open your browser’s DevTools (F12), go to the Network tab, adjust one of the filters on the site, and look for the XHR/fetch request that gets sent. Copy its full Request URL—that’s the endpoint we’ll hit.
2. Write the Code
import requests # Replace this with the actual Ajax endpoint URL from DevTools ajax_endpoint = "https://www.crowdfundmarkt.nl/your-actual-ajax-url-here" # Construct the form data exactly as it appears in the request form_payload = { 'filters[targetValueMin]': '0', 'filters[targetValueMax]': '15000000', 'filters[interestRateMin]': '0', 'filters[interestRateMax]': '20', 'filters[loanTermMin]': '0', 'filters[loanTermMax]': '240', 'filters[platforms]': '', 'filters[financingTypes]': '', 'filters[developmentPhases]': '', 'filters[investmentPurposes]': '' } # Add headers to mimic a browser request (avoids basic anti-scraping checks) request_headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Referer': 'https://www.crowdfundmarkt.nl/crowdfundingprojecten', 'X-Requested-With': 'XMLHttpRequest' # Tells the server it's an Ajax request } # Send the POST request with form data (use data=, not json=!) response = requests.post(ajax_endpoint, data=form_payload, headers=request_headers) # Handle the response if response.status_code == 200: # Most Ajax calls return JSON—parse it scraped_data = response.json() print("Success! Here's a snippet of the data:") print(scraped_data[:5]) # Print first 5 items to verify else: print(f"Oops, request failed. Status code: {response.status_code}") print("Check your URL, headers, or form data for typos!")
Key Tips to Avoid Issues
- Match the keys exactly: Don’t change
filters[targetValueMin]to something likefilters.targetValueMin—the server expects the bracket syntax. - Use a session for cookies: If the site requires you to be "logged in" (even anonymously), use
requests.Session()to persist cookies across requests. - Copy all relevant headers: If the request still fails, copy all headers from the browser’s DevTools request (like
Accept,Accept-Language) and add them to yourrequest_headersdictionary. - Test with
response.textfirst: If you’re not getting JSON, printresponse.textto see if it’s HTML, an error message, or something else.
内容的提问来源于stack exchange,提问作者Ger

