You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬取第二页却返回第一页数据问题求助

Troubleshooting: Scraping Second Page Returns First Page Content

It’s a super common hiccup with dynamic websites—let’s break down the most likely fixes for your issue.

1. Wait for the Page to Fully Load

Dynamic sites often load content asynchronously after the initial page load. Your code grabs the HTML right after browser.get(), which might be too early before the second page’s content finishes rendering.

Fix: Use Selenium’s WebDriverWait to pause until a unique element from page 2 is present (like a page number label or an item exclusive to page 2). Here’s how to adjust your code:

from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By

# Navigate to the second page URL
browser.get("https://XXXXXXXXX/0_9b34?P=2")

# Wait up to 10 seconds for a page-2-specific element to load
# Replace the XPath with something unique to your second page (e.g., a "Page 2" indicator)
wait = WebDriverWait(browser, 10)
wait.until(EC.presence_of_element_located((By.XPATH, "//*[contains(text(), 'Page 2')]")))

# Now extract the fully loaded HTML
innerHTML = browser.execute_script("return document.body.innerHTML")
soup = BeautifulSoup(innerHTML, 'html.parser')
project_items = soup.find_all('td', attrs={'headers': 'ID Item'})

2. Verify the URL Parameter Actually Works

Some sites don’t use URL parameters for pagination—they rely on session cookies, AJAX requests, or JavaScript triggers instead.

Quick Check: Manually open https://XXXXXXXXX/0_9b34?P=2 in your browser. If it still shows page 1, modifying the URL won’t work. Instead, simulate clicking the "Next Page" button:

# Start on the first page
browser.get("https://XXXXXXXXX/0_9b34?P=1")

# Wait for the next button to be clickable, then click it
next_button = wait.until(EC.element_to_be_clickable((By.XPATH, "//a[text()='Next']")))
next_button.click()

# Wait for page 2 content to load before extracting HTML
wait.until(EC.presence_of_element_located((By.XPATH, "//*[contains(text(), 'Page 2')]")))
innerHTML = browser.execute_script("return document.body.innerHTML")

3. Disable Browser Caching

Occasionally, browsers cache old page content, leading to the first page being returned even when navigating to a new URL.

Fix: Configure your Selenium browser to skip caching. For Chrome, add these options:

from selenium.webdriver.chrome.options import Options

chrome_options = Options()
chrome_options.add_argument("--disable-cache")
chrome_options.add_argument("--disable-application-cache")
chrome_options.add_argument("--disable-offline-load-stale-cache")

# Initialize the browser with these settings
browser = webdriver.Chrome(options=chrome_options)

4. Check for Element Attribute Changes

It’s possible the headers attribute for your target td elements is different on page 2. Manually inspect the second page’s HTML to confirm the headers value still matches ID Item. If not, update your find_all selector to match the correct attribute.


内容的提问来源于stack exchange,提问作者Hamza Dellam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:52:21