You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

同一URL呈现不同结果的技术原理及HTTP请求提取IEEE表格数据咨询

IEEE Fellows目录分页实现与数据提取指南

Great question! Let's break this down clearly into two parts: how the pagination trick works, and how to pull the table data without relying on a headless browser.

一、分页实现的具体技术

This "same URL, different table content" behavior is almost always driven by front-end async loading + hash routing (or client-side state management). Here's the step-by-step breakdown:

  • When you click a pagination button (like "2" or "3"), JavaScript intercepts the default link navigation (so the main URL doesn't change).
  • The script updates the URL's hash fragment (e.g., from #results_table to #results_table?page=2 or just #page=2). This keeps the address bar showing the original URL but adds a hidden marker for the current page.
  • The front-end listens for the hashchange event. When it detects a change, it uses AJAX or the Fetch API to send an async request to the backend, passing the current page number as a parameter (like page=2).
  • The backend returns the data for that page (usually as JSON, or sometimes a snippet of HTML), and the front-end JavaScript replaces the existing table content with the new data.

A less likely alternative for smaller datasets is client-side pagination: where all data is loaded upfront into hidden DOM elements or JS variables, and clicking pagination buttons just toggles which subset of data is displayed. But given the size of the IEEE Fellows directory, async loading is the far more probable approach.

二、无需无头浏览器的数据提取方法

To scrape this data without a headless browser, you just need to find and mimic the backend API that feeds the table. Here's how to do it:

1. Inspect the network traffic

  • Open your browser's DevTools (F12 key), go to the Network tab, and filter for "XHR" or "Fetch" requests.
  • Click a pagination button, and watch for a new request appearing in the list. This is the API endpoint that returns the paginated data.
  • Note down critical details: the request URL, HTTP method (GET/POST), any query parameters or form data, and request headers (especially Cookie, User-Agent, and Referer—these are often required to bypass anti-scraping checks).

2. Simulate the request with an HTTP client

You can use tools like curl or programming libraries like Python's requests to replicate the request. Here's a quick Python example:

import requests
from bs4 import BeautifulSoup
import time

# Replace these values with what you captured in DevTools
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
    'Referer': 'https://services27.ieee.org/fellowsdirectory/home.html#results_table',
    'Cookie': 'your-cookie-value-from-devtools'
}

base_url = 'https://services27.ieee.org/your-found-api-endpoint'

# Loop through pages (adjust range based on total pages)
for page_num in range(1, 10):
    params = {
        'page': page_num,
        'size': 20  # Match the page size from the site
    }
    
    response = requests.get(base_url, headers=headers, params=params)
    
    # If response is JSON
    # data = response.json()
    # Process the JSON data here
    
    # If response is HTML snippet
    soup = BeautifulSoup(response.text, 'html.parser')
    rows = soup.find_all('tr')
    for row in rows:
        # Extract data from each table row
        print([cell.text.strip() for cell in row.find_all('td')])
    
    # Add a delay to avoid hitting rate limits
    time.sleep(1)

3. Handle pagination fully

  • Check the first response for metadata like total pages or total items (often included in JSON responses). Use this to set the upper limit for your loop instead of hardcoding a number.
  • If the site uses POST requests instead of GET, move the parameters into the data argument of the requests.post() call.

4. Key notes

  • Anti-scraping measures: Some sites block frequent requests, so add delays between requests. If your cookie expires, re-capture it from DevTools.
  • Authentication: IEEE's services may require a valid session, so you might need to first send a request to the main page to get a fresh cookie before making API calls.

内容的提问来源于stack exchange,提问作者user1424739

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:17:55