使用Selenium提取网站多页数据失败,请求技术排查指导
代码错误分析及修正
核心错误点
未更新页面解析对象
你只在程序启动时获取了一次页面源码并创建BeautifulSoup对象,翻页后页面内容已经更新,但始终用第一次的soup解析,自然只能拿到第一页数据。同时if(count==0)的判断直接限制了只有第一次循环会提取数据,后续翻页后完全没执行提取逻辑。未定义变量
main_url
函数里写了url=main_url,但整个代码里都没定义过这个变量,运行时会直接报错。翻页后无可靠加载验证
用固定的time.sleep(5)不够稳妥,应该等待页面上的产品元素更新后再提取,避免因加载慢导致的数据抓取不完整。
修正后的代码
from bs4 import BeautifulSoup import pandas as pd from urllib3.exceptions import InsecureRequestWarning from urllib3 import disable_warnings import time from selenium import webdriver from selenium.webdriver.chrome.service import Service from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # 禁用不安全请求警告 disable_warnings(InsecureRequestWarning) url = 'https://cargillsonline.com/Web/Product?IC=Mg==&NC=QmFieSBQcm9kdWN0cw==' path = 'C:/Users/dell/Desktop/Data/DataScraping/chrome_driver/chromedriver' service = Service(path) driver = webdriver.Chrome(service=service) driver.get(url) def get_data(): start = time.process_time() product_name = [] product_price = [] all_pages = 10 # 测试用页数 print('Get Data Processing .....') for i in range(all_pages): # 每次循环都重新获取当前页面的源码并创建新的soup对象 html = driver.page_source soup = BeautifulSoup(html, 'html.parser') # 提取当前页的产品名称 add_boxs_v1 = soup.find_all(class_='veg') for product in add_boxs_v1: product_name.append(product.find('p').text.strip()) # 提取当前页的产品价格 add_boxs_v2 = soup.find_all(class_='strike1') for price in add_boxs_v2: product_price.append(price.find('h4').text.strip()) # 不是最后一页的话,执行翻页操作 if i < all_pages - 1: try: # 等待下一页按钮可点击并点击 next_btn = WebDriverWait(driver, 20).until( EC.element_to_be_clickable((By.XPATH, "//a[@ng-click='selectPage(page + 1, $event)']")) ) next_btn.click() # 等待产品元素更新,确保页面加载完成 WebDriverWait(driver, 20).until( EC.staleness_of(add_boxs_v1[0]) if add_boxs_v1 else EC.presence_of_element_located((By.CLASS_NAME, 'veg')) ) except Exception as e: print(f"翻页失败: {e}") break print('done') print(f"耗时: {time.process_time() - start:.2f}秒") df = pd.DataFrame({'Product_name': product_name, 'Price': product_price}) return df df = get_data() print(df.head()) # 关闭浏览器 driver.quit()
内容的提问来源于stack exchange,提问作者Chanaka Eranga
相关产品推荐
相关产品推荐

