Selenium爬取国家兽药动态网站返回None,请求技术指导
Got it, let's work through this problem together. The site you're targeting is a classic single-page application (SPA)—it loads content dynamically via backend API calls instead of updating the URL when you click sidebar items. That's why your initial Selenium setup isn't returning data directly; the actual JSON payloads are coming from hidden API requests, not the main page HTML.
Here's a step-by-step guide to fix this:
1. First, Find the Real API Endpoints via DevTools
This is the most critical step. You need to uncover the actual requests the site makes to fetch data:
- Open Chrome (or any Chromium-based browser) and navigate to the site.
- Hit
F12to open DevTools, then switch to the Network tab. - Check the Preserve log checkbox at the top of the Network tab, and filter requests by XHR/Fetch (this will show only the dynamic data requests).
- Now click any sidebar item you want to scrape. You'll see new requests pop up in the Network list. Look for the one that returns a JSON response (click on a request, go to the Response tab to check).
- Note down:
- The full request URL (it might be something like
http://124.126.15.169:8081/api/xxx) - The request method (GET or POST)
- Any query parameters or form data (check the Params or Payload tab)
- Important request headers (like
Cookie,User-Agent, orReferer—these are often required to avoid being blocked)
- The full request URL (it might be something like
2. Option 1: Use Selenium to Capture API Responses
If you need to keep using Selenium (e.g., to handle login or session state), you can listen for network requests and extract the JSON data directly:
Here's a quick Python example:
from selenium import webdriver from selenium.webdriver.chrome.options import Options import json # Set up Chrome to log network performance data chrome_options = Options() chrome_options.set_capability("goog:loggingPrefs", {"performance": "ALL"}) driver = webdriver.Chrome(options=chrome_options) driver.get("http://124.126.15.169:8081/cx/") # Click your target sidebar item here (replace with your actual selector) driver.find_element("xpath", "//div[text()='某个栏目']").click() # Parse the performance logs to find the API request logs = driver.get_log("performance") for log in logs: log_json = json.loads(log["message"])["message"] # Look for the response event of your target API if log_json["method"] == "Network.responseReceived" and "json" in log_json["params"]["response"]["mimeType"]: request_id = log_json["params"]["requestId"] # Get the actual response body response_body = driver.execute_cdp_cmd("Network.getResponseBody", {"requestId": request_id}) print(json.loads(response_body["body"])) break driver.quit()
3. Option 2: Call the API Directly (Faster & More Efficient)
Once you have the API details from DevTools, you can skip Selenium entirely and use requests to fetch the JSON data directly. This is way faster than driving a browser:
import requests # Replace with the actual API URL, headers, and params you found api_url = "http://124.126.15.169:8081/xxx/xxx" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36", "Cookie": "your_cookie_here", # Get this from DevTools or Selenium "Referer": "http://124.126.15.169:8081/cx/" } params = { "type": "xxx", # Replace with actual parameters from DevTools "page": 1 } response = requests.get(api_url, headers=headers, params=params) data = response.json() print(data)
Note: If the site requires a session (like after logging in), you can use requests.Session() to maintain cookies across requests, or grab the initial cookies from Selenium and pass them to requests.
4. Common Pitfalls to Watch For
- Encrypted Parameters: Some sites encrypt request parameters (like
signortoken). If you see this, you'll need to analyze the site's frontend JavaScript to reverse-engineer the encryption logic. - Session Expiry: Cookies might expire after a while—if your requests start failing, refresh the cookies via DevTools or Selenium.
- Anti-Scraping: The site might block repeated requests. Add delays between requests, use a rotating User-Agent, or avoid hitting the API too hard.
内容的提问来源于stack exchange,提问作者riversxiao

