You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium提取网站多页数据失败,请求技术排查指导

代码错误分析及修正

核心错误点

  • 未更新页面解析对象
    你只在程序启动时获取了一次页面源码并创建BeautifulSoup对象,翻页后页面内容已经更新,但始终用第一次的soup解析,自然只能拿到第一页数据。同时if(count==0)的判断直接限制了只有第一次循环会提取数据,后续翻页后完全没执行提取逻辑。

  • 未定义变量main_url
    函数里写了url=main_url,但整个代码里都没定义过这个变量,运行时会直接报错。

  • 翻页后无可靠加载验证
    用固定的time.sleep(5)不够稳妥,应该等待页面上的产品元素更新后再提取,避免因加载慢导致的数据抓取不完整。


修正后的代码

from bs4 import BeautifulSoup
import pandas as pd
from urllib3.exceptions import InsecureRequestWarning
from urllib3 import disable_warnings
import time
from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

# 禁用不安全请求警告
disable_warnings(InsecureRequestWarning)

url = 'https://cargillsonline.com/Web/Product?IC=Mg==&NC=QmFieSBQcm9kdWN0cw=='
path = 'C:/Users/dell/Desktop/Data/DataScraping/chrome_driver/chromedriver'

service = Service(path)  
driver = webdriver.Chrome(service=service)
driver.get(url)

def get_data():
    start = time.process_time()
    product_name = []
    product_price = []
    all_pages = 10  # 测试用页数
    print('Get Data Processing .....')

    for i in range(all_pages):
        # 每次循环都重新获取当前页面的源码并创建新的soup对象
        html = driver.page_source
        soup = BeautifulSoup(html, 'html.parser')
        
        # 提取当前页的产品名称
        add_boxs_v1 = soup.find_all(class_='veg')
        for product in add_boxs_v1:
            product_name.append(product.find('p').text.strip())
        
        # 提取当前页的产品价格
        add_boxs_v2 = soup.find_all(class_='strike1')
        for price in add_boxs_v2:
            product_price.append(price.find('h4').text.strip())
        
        # 不是最后一页的话,执行翻页操作
        if i < all_pages - 1:
            try:
                # 等待下一页按钮可点击并点击
                next_btn = WebDriverWait(driver, 20).until(
                    EC.element_to_be_clickable((By.XPATH, "//a[@ng-click='selectPage(page + 1, $event)']"))
                )
                next_btn.click()
                # 等待产品元素更新,确保页面加载完成
                WebDriverWait(driver, 20).until(
                    EC.staleness_of(add_boxs_v1[0]) if add_boxs_v1 else EC.presence_of_element_located((By.CLASS_NAME, 'veg'))
                )
            except Exception as e:
                print(f"翻页失败: {e}")
                break

    print('done')
    print(f"耗时: {time.process_time() - start:.2f}秒")
    df = pd.DataFrame({'Product_name': product_name, 'Price': product_price})
    return df

df = get_data()
print(df.head())
# 关闭浏览器
driver.quit()

内容的提问来源于stack exchange,提问作者Chanaka Eranga

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 16:10:41