如何解决Python Selenium网页爬虫函数中的IndexError问题?
Selenium爬取多类别页面时IndexError问题的解决方案
问题背景
使用Selenium爬取Aldi英国官网的产品信息,单类别页面(含40个产品)爬取正常,但循环处理CSV中的多类别URL时,反复触发IndexError: list index out of range错误。因特定限制无法使用Scrapy-Playwright,需修复当前Selenium代码。
错误追踪信息
Traceback (most recent call last): File "h:\Python\aldi\aldi\selenium_test2.py", line 76, in <module> scraper(cat_urls) File "h:\Python\aldi\aldi\selenium_test2.py", line 71, in scraper temp_prod_info = [ {'product_id': id_list[i], 'product_name': name_list[i], 'price': price_list[i] } for i in range(len(names)) ] File "h:\Python\aldi\aldi\selenium_test2.py", line 71, in <listcomp> temp_prod_info = [ {'product_id': id_list[i], 'product_name': name_list[i], 'price': price_list[i] } for i in range(len(names)) ] IndexError: list index out of range
CSV文件内容(url_list.csv)
https://groceries.aldi.co.uk/en-GB/bakery https://groceries.aldi.co.uk/en-GB/fresh-food
原始代码
# from weakref import proxy from selenium import webdriver from selenium.webdriver.common.keys import Keys import pandas as pd from selenium.webdriver.common.by import By import time from fp.fp import FreeProxy options = webdriver.ChromeOptions() options.add_experimental_option("excludeSwitches", ["enable-logging"]) options.add_argument('--disable-blink-features=AutomationControlled') with open('url_list.csv') as f: cat_urls = [line.strip() for line in f] name_list=[] price_list=[] href_list=[] id_list= [] prod_info= [] proxy = FreeProxy(rand=True, country_id=['GB']).get() def scraper(cat_urls): for cat_url in cat_urls: options.add_argument('--proxy-server=%s' % proxy) driver = webdriver.Chrome(options=options) driver.get(cat_url) driver.find_element('xpath','//*[@id="onetrust-accept-btn-handler"]').click() time.sleep(5) names = driver.find_elements(By.CLASS_NAME, 'product-tile-text.text-center.px-3.mb-3') def name_lister (names, name_list): for name in names: text_name = name.text name_list.append(text_name) return name_list name_lister(names, name_list) prices = driver.find_elements(By.CSS_SELECTOR, 'span.h4') def price_lister (prices, price_list): for price in prices: text_price = price.text.strip('£') price_list.append(text_price) return price_list price_lister(prices, price_list) hrefs = driver.find_elements(By.CLASS_NAME, 'p.text-default-font') def href_lister (hrefs, href_list): for href in hrefs: text_id = href.get_attribute('href') href_list.append(text_id) return href_list href_lister(hrefs, href_list) def id_splitter (href_list, id_list): for href in href_list: if href is not None: id = href[-13:] else: id = "" id_list.append(id) return id_list id_splitter(href_list, id_list) # prod_info = [ {'product_id': id_list[i], 'product_name': name_list[i], 'price': price_list[i] } for i in range(len(name_list)) ] temp_prod_info = [ {'product_id': id_list[i], 'product_name': name_list[i], 'price': price_list[i] } for i in range(len(names)) ] prod_info.append(temp_prod_info) return prod_info scraper(cat_urls) df = pd.DataFrame.from_dict(prod_info, orient='columns') df.to_csv("scrape_data.csv")
问题核心原因
- 全局列表累积导致索引不匹配:
name_list、price_list、id_list为全局变量,每爬取一个类别就追加数据。处理第二个类别时,len(names)是当前页面产品数,但全局列表已包含第一个类别的数据,导致索引超出范围。 id_splitter逻辑错误:仅当href为None时才往id_list追加空值,正常id未被添加,导致id_list长度远小于其他列表,直接触发索引越界。- 重复添加代理配置:每次循环都给
options添加代理参数,导致代理重复设置,可能引发浏览器异常。 - 未关闭浏览器进程:爬取完页面后未关闭
driver,残留大量进程影响性能。
修复后的代码
from selenium import webdriver from selenium.webdriver.common.by import By import pandas as pd import time from fp.fp import FreeProxy def scraper(cat_urls): prod_info = [] proxy = FreeProxy(rand=True, country_id=['GB']).get() for cat_url in cat_urls: # 每个页面重新初始化配置,避免重复添加代理 options = webdriver.ChromeOptions() options.add_experimental_option("excludeSwitches", ["enable-logging"]) options.add_argument('--disable-blink-features=AutomationControlled') options.add_argument('--proxy-server=%s' % proxy) driver = webdriver.Chrome(options=options) try: driver.get(cat_url) # 处理Cookie弹窗,添加等待避免元素未加载 time.sleep(3) accept_btn = driver.find_element(By.ID, 'onetrust-accept-btn-handler') if accept_btn.is_displayed(): accept_btn.click() time.sleep(5) # 定位当前页面所有产品父元素,确保信息属于同一产品 products = driver.find_elements(By.CLASS_NAME, 'product-tile') page_prod_info = [] for product in products: # 从单个产品元素内提取信息,彻底避免列表不匹配 name_elem = product.find_element(By.CLASS_NAME, 'product-tile-text.text-center.px-3.mb-3') product_name = name_elem.text price_elem = product.find_element(By.CSS_SELECTOR, 'span.h4') price = price_elem.text.strip('£') # 修正字符乱码问题 href_elem = product.find_element(By.CLASS_NAME, 'p.text-default-font') href = href_elem.get_attribute('href') product_id = href[-13:] if href else "" page_prod_info.append({ 'product_id': product_id, 'product_name': product_name, 'price': price }) prod_info.extend(page_prod_info) finally: # 无论是否出错都关闭浏览器,释放资源 driver.quit() return prod_info # 读取URL列表 with open('url_list.csv') as f: cat_urls = [line.strip() for line in f] # 执行爬取并保存数据 prod_data = scraper(cat_urls) df = pd.DataFrame(prod_data) df.to_csv("scrape_data.csv", index=False)
关键修复点
- 摒弃全局列表:每个类别页面单独处理数据,直接生成当前页面的产品字典列表,避免累积导致的索引混乱。
- 基于单产品元素提取信息:先定位每个产品的父元素,再从父元素内提取名称、价格、链接,确保三个信息属于同一产品,彻底解决列表长度不匹配问题。
- 修正ID提取逻辑:正常情况下的
product_id也会被添加,不再遗漏数据。 - 循环内初始化浏览器配置:每次爬取新页面时重新创建
options,避免重复添加代理参数。 - 强制关闭浏览器:使用
finally块确保每个页面爬取完成后关闭浏览器,释放系统资源。 - 修复价格字符问题:直接用
£替换乱码的£,保证价格数据准确。
内容的提问来源于stack exchange,提问作者Chris
相关产品推荐
相关产品推荐

