You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

分页Web Scraping问题求助:仅爬取第一页数据如何解决?

多页爬取Tenet Diagnostics测试数据失败问题排查与修复

我正在使用Web Scraping从Tenet Diagnostics网站爬取测试数据,但卡在了多页数据提取环节。检查分页源码后,代码仍仅返回第一页数据,以下是我使用的代码:

import csv
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

def navigate_to_next_page():
    try:
        next_button = WebDriverWait(driver, 60).until(
            EC.element_to_be_clickable((By.CSS_SELECTOR, 'li.pagination-next a'))
        )
        next_button.click()
        return True
    except:
        return False

def extract_test_data():
    # Find all test divs
    test_divs = WebDriverWait(driver, 60).until(
        EC.presence_of_all_elements_located((By.CSS_SELECTOR, "div.product"))
    )
    # Iterate over each test div to extract test name, URL, and price
    for test_div in test_divs:
        test_link = test_div.find_element(By.CSS_SELECTOR, "a.text-theme-colored")
        test_name = test_link.text.strip()
        test_url = test_link.get_attribute("href")  # Extract href attribute for URL
        test_price = test_div.find_element(By.CSS_SELECTOR, "span.amount").text.strip()
        # Append test data to the list
        all_test_data.append([test_url, test_name, test_price])

base_url = "https://www.tenetdiagnostics.in/book/tests?type=p"

chrome_options = Options()
chrome_options.add_argument("--headless")
driver = webdriver.Chrome(options=chrome_options)

driver.get(base_url)

all_test_data = []

while True:
    extract_test_data()
    if not navigate_to_next_page():
        break

csv_file = "tenet_test_data.csv"
with open(csv_file, "w", newline="", encoding="utf-8") as file:
    writer = csv.writer(file)
    writer.writerow(["Test URL", "Test Name", "Test Price"])  # Write header
    writer.writerows(all_test_data)

print("Test data saved to", csv_file)

driver.quit()

该代码能返回第一页的预期结果,但我需要提取所有页面的数据。


问题分析与修复方案

1. 分页按钮状态判断缺失

原代码仅通过element_to_be_clickable定位按钮,但到达最后一页时,按钮可能仍存在但被标记为禁用(比如添加disabled类),此时点击操作无效但函数仍返回True,导致循环无法终止。需要先检查按钮的禁用状态。

2. 页面切换后无刷新等待

点击下一页后,页面需要时间加载新数据,原代码没有等待页面完全刷新,可能重复提取第一页的缓存元素。需添加等待逻辑,确保页面切换完成。

3. 无头模式布局异常

无头模式下Chrome的渲染窗口默认较小,可能导致页面元素布局错位,元素定位失败。需设置窗口大小适配页面。


修改后的完整代码

import csv
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.common.exceptions import TimeoutException, StaleElementReferenceException

def navigate_to_next_page():
    try:
        # 先定位分页按钮的父容器,检查是否禁用
        next_li = WebDriverWait(driver, 30).until(
            EC.presence_of_element_located((By.CSS_SELECTOR, 'li.pagination-next'))
        )
        if 'disabled' in next_li.get_attribute('class'):
            return False
        
        next_button = next_li.find_element(By.TAG_NAME, 'a')
        WebDriverWait(driver, 30).until(EC.element_to_be_clickable(next_button))
        
        # 记录当前页最后一个测试链接,用于验证页面是否切换
        last_test_url = all_test_data[-1][0] if all_test_data else ""
        next_button.click()
        
        # 等待页面切换:要么URL变化,要么出现新的测试数据
        WebDriverWait(driver, 30).until(
            lambda d: (len(d.find_elements(By.CSS_SELECTOR, "div.product")) > 0 and 
                      d.find_elements(By.CSS_SELECTOR, "div.product a.text-theme-colored")[0].get_attribute("href") != last_test_url) or
                      ('page=' in d.current_url and d.current_url != base_url)
        )
        return True
    except (TimeoutException, StaleElementReferenceException):
        return False

def extract_test_data():
    try:
        test_divs = WebDriverWait(driver, 30).until(
            EC.presence_of_all_elements_located((By.CSS_SELECTOR, "div.product"))
        )
        # 避免重复添加同一页面数据
        for test_div in test_divs:
            test_link = test_div.find_element(By.CSS_SELECTOR, "a.text-theme-colored")
            test_url = test_link.get_attribute("href")
            if not any(item[0] == test_url for item in all_test_data):
                test_name = test_link.text.strip()
                test_price = test_div.find_element(By.CSS_SELECTOR, "span.amount").text.strip()
                all_test_data.append([test_url, test_name, test_price])
    except StaleElementReferenceException:
        # 元素过期时重新提取
        extract_test_data()

base_url = "https://www.tenetdiagnostics.in/book/tests?type=p"

chrome_options = Options()
chrome_options.add_argument("--headless")
# 设置窗口大小,避免无头模式布局异常
chrome_options.add_argument("--window-size=1920,1080")
driver = webdriver.Chrome(options=chrome_options)

driver.get(base_url)

all_test_data = []

while True:
    extract_test_data()
    if not navigate_to_next_page():
        break

csv_file = "tenet_test_data.csv"
with open(csv_file, "w", newline="", encoding="utf-8") as file:
    writer = csv.writer(file)
    writer.writerow(["Test URL", "Test Name", "Test Price"])
    writer.writerows(all_test_data)

print("Test data saved to", csv_file)

driver.quit()

内容的提问来源于stack exchange,提问作者Diksha Ingle

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 05:20:33