同一URL呈现不同结果的技术原理及HTTP请求提取IEEE表格数据咨询
Great question! Let's break this down clearly into two parts: how the pagination trick works, and how to pull the table data without relying on a headless browser.
一、分页实现的具体技术
This "same URL, different table content" behavior is almost always driven by front-end async loading + hash routing (or client-side state management). Here's the step-by-step breakdown:
- When you click a pagination button (like "2" or "3"), JavaScript intercepts the default link navigation (so the main URL doesn't change).
- The script updates the URL's hash fragment (e.g., from
#results_tableto#results_table?page=2or just#page=2). This keeps the address bar showing the original URL but adds a hidden marker for the current page. - The front-end listens for the
hashchangeevent. When it detects a change, it uses AJAX or the Fetch API to send an async request to the backend, passing the current page number as a parameter (likepage=2). - The backend returns the data for that page (usually as JSON, or sometimes a snippet of HTML), and the front-end JavaScript replaces the existing table content with the new data.
A less likely alternative for smaller datasets is client-side pagination: where all data is loaded upfront into hidden DOM elements or JS variables, and clicking pagination buttons just toggles which subset of data is displayed. But given the size of the IEEE Fellows directory, async loading is the far more probable approach.
二、无需无头浏览器的数据提取方法
To scrape this data without a headless browser, you just need to find and mimic the backend API that feeds the table. Here's how to do it:
1. Inspect the network traffic
- Open your browser's DevTools (F12 key), go to the Network tab, and filter for "XHR" or "Fetch" requests.
- Click a pagination button, and watch for a new request appearing in the list. This is the API endpoint that returns the paginated data.
- Note down critical details: the request URL, HTTP method (GET/POST), any query parameters or form data, and request headers (especially
Cookie,User-Agent, andReferer—these are often required to bypass anti-scraping checks).
2. Simulate the request with an HTTP client
You can use tools like curl or programming libraries like Python's requests to replicate the request. Here's a quick Python example:
import requests from bs4 import BeautifulSoup import time # Replace these values with what you captured in DevTools headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Referer': 'https://services27.ieee.org/fellowsdirectory/home.html#results_table', 'Cookie': 'your-cookie-value-from-devtools' } base_url = 'https://services27.ieee.org/your-found-api-endpoint' # Loop through pages (adjust range based on total pages) for page_num in range(1, 10): params = { 'page': page_num, 'size': 20 # Match the page size from the site } response = requests.get(base_url, headers=headers, params=params) # If response is JSON # data = response.json() # Process the JSON data here # If response is HTML snippet soup = BeautifulSoup(response.text, 'html.parser') rows = soup.find_all('tr') for row in rows: # Extract data from each table row print([cell.text.strip() for cell in row.find_all('td')]) # Add a delay to avoid hitting rate limits time.sleep(1)
3. Handle pagination fully
- Check the first response for metadata like total pages or total items (often included in JSON responses). Use this to set the upper limit for your loop instead of hardcoding a number.
- If the site uses POST requests instead of GET, move the parameters into the
dataargument of therequests.post()call.
4. Key notes
- Anti-scraping measures: Some sites block frequent requests, so add delays between requests. If your cookie expires, re-capture it from DevTools.
- Authentication: IEEE's services may require a valid session, so you might need to first send a request to the main page to get a fresh cookie before making API calls.
内容的提问来源于stack exchange,提问作者user1424739

