You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解决Python Selenium网页爬虫函数中的IndexError问题?

Selenium爬取多类别页面时IndexError问题的解决方案

问题背景

使用Selenium爬取Aldi英国官网的产品信息,单类别页面(含40个产品)爬取正常,但循环处理CSV中的多类别URL时,反复触发IndexError: list index out of range错误。因特定限制无法使用Scrapy-Playwright,需修复当前Selenium代码。

错误追踪信息

Traceback (most recent call last):
  File "h:\Python\aldi\aldi\selenium_test2.py", line 76, in <module>
    scraper(cat_urls)
  File "h:\Python\aldi\aldi\selenium_test2.py", line 71, in scraper
    temp_prod_info = [ {'product_id': id_list[i], 'product_name': name_list[i], 'price': price_list[i] } for i in range(len(names)) ]
  File "h:\Python\aldi\aldi\selenium_test2.py", line 71, in <listcomp>
    temp_prod_info = [ {'product_id': id_list[i], 'product_name': name_list[i], 'price': price_list[i] } for i in range(len(names)) ]
IndexError: list index out of range

CSV文件内容(url_list.csv)

https://groceries.aldi.co.uk/en-GB/bakery
https://groceries.aldi.co.uk/en-GB/fresh-food

原始代码

# from weakref import proxy
from selenium import webdriver
from selenium.webdriver.common.keys import Keys
import pandas as pd
from selenium.webdriver.common.by import By
import time
from fp.fp import FreeProxy


options = webdriver.ChromeOptions() 
options.add_experimental_option("excludeSwitches", ["enable-logging"])
options.add_argument('--disable-blink-features=AutomationControlled')

with open('url_list.csv') as f:
    cat_urls = [line.strip() for line in f]

name_list=[]
price_list=[]
href_list=[]
id_list= []
prod_info= []

proxy = FreeProxy(rand=True, country_id=['GB']).get()

def scraper(cat_urls):
    for cat_url in cat_urls:
        options.add_argument('--proxy-server=%s' % proxy)
        driver = webdriver.Chrome(options=options)
        driver.get(cat_url)
        driver.find_element('xpath','//*[@id="onetrust-accept-btn-handler"]').click()
        time.sleep(5)
        
        names = driver.find_elements(By.CLASS_NAME, 'product-tile-text.text-center.px-3.mb-3')
               
        def name_lister (names, name_list):
            for name in names:
                text_name = name.text
                name_list.append(text_name)
            return name_list
        name_lister(names, name_list)

        prices = driver.find_elements(By.CSS_SELECTOR, 'span.h4')
        
        def price_lister (prices, price_list):
            for price in prices:
                text_price = price.text.strip('£')
                price_list.append(text_price)
            return price_list
        price_lister(prices, price_list)

        hrefs = driver.find_elements(By.CLASS_NAME, 'p.text-default-font')
        
        def href_lister (hrefs, href_list):
            for href in hrefs:
                text_id = href.get_attribute('href')
                href_list.append(text_id)
            return href_list
        href_lister(hrefs, href_list)
        
        def id_splitter (href_list, id_list):
            for href in href_list:
                if href is not None:
                    id = href[-13:]
                else:
                    id = ""
                    id_list.append(id)
            return id_list
        id_splitter(href_list, id_list)
        
        # prod_info = [ {'product_id': id_list[i], 'product_name': name_list[i], 'price': price_list[i] } for i in range(len(name_list)) ]
        temp_prod_info = [ {'product_id': id_list[i], 'product_name': name_list[i], 'price': price_list[i] } for i in range(len(names)) ]
        prod_info.append(temp_prod_info)
      
    return prod_info

scraper(cat_urls)

df = pd.DataFrame.from_dict(prod_info, orient='columns')
df.to_csv("scrape_data.csv")

问题核心原因

  1. 全局列表累积导致索引不匹配:name_list、price_list、id_list为全局变量,每爬取一个类别就追加数据。处理第二个类别时,len(names)是当前页面产品数,但全局列表已包含第一个类别的数据,导致索引超出范围。
  2. id_splitter逻辑错误:仅当href为None时才往id_list追加空值,正常id未被添加,导致id_list长度远小于其他列表,直接触发索引越界。
  3. 重复添加代理配置:每次循环都给options添加代理参数,导致代理重复设置,可能引发浏览器异常。
  4. 未关闭浏览器进程:爬取完页面后未关闭driver,残留大量进程影响性能。

修复后的代码

from selenium import webdriver
from selenium.webdriver.common.by import By
import pandas as pd
import time
from fp.fp import FreeProxy

def scraper(cat_urls):
    prod_info = []
    proxy = FreeProxy(rand=True, country_id=['GB']).get()
    
    for cat_url in cat_urls:
        # 每个页面重新初始化配置,避免重复添加代理
        options = webdriver.ChromeOptions() 
        options.add_experimental_option("excludeSwitches", ["enable-logging"])
        options.add_argument('--disable-blink-features=AutomationControlled')
        options.add_argument('--proxy-server=%s' % proxy)
        
        driver = webdriver.Chrome(options=options)
        try:
            driver.get(cat_url)
            # 处理Cookie弹窗,添加等待避免元素未加载
            time.sleep(3)
            accept_btn = driver.find_element(By.ID, 'onetrust-accept-btn-handler')
            if accept_btn.is_displayed():
                accept_btn.click()
            time.sleep(5)
            
            # 定位当前页面所有产品父元素,确保信息属于同一产品
            products = driver.find_elements(By.CLASS_NAME, 'product-tile')
            
            page_prod_info = []
            for product in products:
                # 从单个产品元素内提取信息,彻底避免列表不匹配
                name_elem = product.find_element(By.CLASS_NAME, 'product-tile-text.text-center.px-3.mb-3')
                product_name = name_elem.text
                
                price_elem = product.find_element(By.CSS_SELECTOR, 'span.h4')
                price = price_elem.text.strip('£')  # 修正字符乱码问题
                
                href_elem = product.find_element(By.CLASS_NAME, 'p.text-default-font')
                href = href_elem.get_attribute('href')
                product_id = href[-13:] if href else ""
                
                page_prod_info.append({
                    'product_id': product_id,
                    'product_name': product_name,
                    'price': price
                })
            
            prod_info.extend(page_prod_info)
        finally:
            # 无论是否出错都关闭浏览器,释放资源
            driver.quit()
    
    return prod_info

# 读取URL列表
with open('url_list.csv') as f:
    cat_urls = [line.strip() for line in f]

# 执行爬取并保存数据
prod_data = scraper(cat_urls)
df = pd.DataFrame(prod_data)
df.to_csv("scrape_data.csv", index=False)

关键修复点

  • 摒弃全局列表:每个类别页面单独处理数据,直接生成当前页面的产品字典列表,避免累积导致的索引混乱。
  • 基于单产品元素提取信息:先定位每个产品的父元素,再从父元素内提取名称、价格、链接,确保三个信息属于同一产品,彻底解决列表长度不匹配问题。
  • 修正ID提取逻辑:正常情况下的product_id也会被添加,不再遗漏数据。
  • 循环内初始化浏览器配置:每次爬取新页面时重新创建options,避免重复添加代理参数。
  • 强制关闭浏览器:使用finally块确保每个页面爬取完成后关闭浏览器,释放系统资源。
  • 修复价格字符问题:直接用£替换乱码的£,保证价格数据准确。

内容的提问来源于stack exchange,提问作者Chris

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 03:40:34