You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium爬取多页房产数据触发IndexError问题求助

Selenium爬虫IndexError问题排查

问题场景

以下是爬取佛罗里达州顶级房产经纪人数据的代码:

from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.common.by import By


elemental_list = []

chrome_driver = Service(executable_path="Users\David\Desktop\Python\chromedriver_win32\chromedriver.exe")
driver = webdriver.Chrome(service=chrome_driver)

for page in range(21):

    page_url = "https://www.fastexpert.com/top-real-estate-agents/florida/?page=" + str(page)
    driver.get(page_url)
    title = driver.find_elements(By.XPATH, "//h3/a")
    location = driver.find_elements(By.XPATH, "//div[contains(@class, 'RTLOCATION')]/p/a")

    for i in range(len(title)):
        elemental_list.append([title[i].text, location[i].text])

for element in elemental_list:
    print(element)

driver.close()

仅循环2次时能正常返回结果,但遍历全部20页时触发索引越界错误:

Traceback (most recent call last):
  File "C:\Users\David\PycharmProjects\pythonProject\main.py", line 19, in <module>
    elemental_list.append([title[i].text, location[i].text])
IndexError: list index out of range

核心原因

  1. 元素加载不同步:部分页面加载速度慢,代码执行元素查找时,title和location的元素未完全加载,导致其中一个列表长度短于另一个,循环时出现索引超出范围。
  2. 页面结构不一致:少数页面的DOM结构与其他页面有差异,存在没有对应location的经纪人条目,使得location列表长度小于title。
  3. 反爬机制干扰:高频请求触发网站反爬限制,部分页面被拦截或内容加载不全,导致元素抓取数量不匹配。

解决办法

1. 增加等待机制,确保元素加载完成

使用显式等待+固定等待组合,保证页面元素加载完毕后再执行查找:

from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import time

elemental_list = []

chrome_driver = Service(executable_path="Users\David\Desktop\Python\chromedriver_win32\chromedriver.exe")
driver = webdriver.Chrome(service=chrome_driver)
wait = WebDriverWait(driver, 10)  # 最长等待10秒

for page in range(21):
    page_url = f"https://www.fastexpert.com/top-real-estate-agents/florida/?page={page}"
    driver.get(page_url)
    time.sleep(1)  # 缓冲等待,避免请求过于频繁
    
    # 显式等待标题元素全部加载
    wait.until(EC.presence_of_all_elements_located((By.XPATH, "//h3/a")))
    title = driver.find_elements(By.XPATH, "//h3/a")
    location = driver.find_elements(By.XPATH, "//div[contains(@class, 'RTLOCATION')]/p/a")
    
    # 取两个列表的最小长度,避免索引越界
    min_length = min(len(title), len(location))
    for i in range(min_length):
        elemental_list.append([title[i].text, location[i].text])

for element in elemental_list:
    print(element)

driver.close()

2. 添加异常捕获,跳过异常条目

在循环中加入异常处理,遇到索引错误时跳过当前条目,不中断整个爬虫流程:

# 保留原有代码结构,修改循环部分
for i in range(len(title)):
    try:
        elemental_list.append([title[i].text, location[i].text])
    except IndexError:
        print(f"第{page}页第{i}条数据无对应位置信息,跳过")
        continue

3. 优化请求频率,规避反爬

使用随机等待模拟人类浏览行为,降低被反爬拦截的概率:

import random
# 在页面请求后添加随机等待
for page in range(21):
    driver.get(page_url)
    time.sleep(random.uniform(1, 3))  # 随机等待1-3秒
    # 后续元素查找代码不变

内容的提问来源于stack exchange,提问作者DaveMier88

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 01:12:15