You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助解决Python爬虫中的IndexError: list index out of range错误

问题排查:爬虫脚本IndexError报错解决

报错原因

IndexError: list index out of range 说明你的代码中property、duration或price这三个变量里,至少有一个是空列表——也就是对应的XPath表达式在当前爬取的页面中没有找到目标元素。可能的诱因包括:

  • 目标页面的HTML结构发生了变化,原XPath路径失效
  • 请求返回的内容异常(比如404页面、反爬拦截页面)
  • XPath写法过于依赖固定的DOM层级,容错性差

修复方案

1. 增加空值检查与错误记录

在尝试获取元素文本前,先判断列表是否为空,避免直接索引空列表;同时记录出错的URL,方便后续针对性排查。

2. 优化XPath表达式

尽量避免使用依赖固定层级的绝对路径,改用元素的特征(比如class属性、文本内容)来定位,提升鲁棒性。例如:

  • 原//*[@id="header"]/div/div[2]/h1 可改为//h1[contains(@class, 'property-title')](假设标题有对应的class)
  • 原//*[@id="price"]/div/div/span/span[3] 可改为//span[contains(@class, 'price-value')](假设价格元素有标识class)

3. 增加请求异常处理

捕获请求过程中可能出现的错误(比如网络超时、连接失败),避免脚本直接崩溃。

4. 优化文件写入逻辑

不要在循环内反复打开/关闭输出文件,改为在循环外打开,提升运行效率。

修改后的完整代码

import requests
from bs4 import BeautifulSoup
from lxml import etree
import csv

# 提前打开输出文件,避免循环内重复IO操作
with open('1_colonia.csv', 'r', encoding='utf-8') as infile, \
     open('2_colonia.csv', 'a', newline='', encoding='utf-8') as outfile:
    reader = csv.reader(infile, delimiter=';')
    writer = csv.writer(outfile, delimiter=';')
    # 可选:如果输出文件是新的,先写入表头
    # writer.writerow(["URL", "Property", "Duration", "Price"])
    
    next(reader)  # 跳过表头
    for row in reader:
        url = row[0]
        try:
            # 增加请求超时与异常捕获
            page = requests.get(url, timeout=10)
            page.raise_for_status()  # 检查请求是否成功(比如404、500会抛出异常)
            
            soup = BeautifulSoup(page.content, 'html.parser')
            dom = etree.HTML(str(soup))
            
            # 用更鲁棒的XPath,同时处理空列表情况
            property_elem = dom.xpath('//*[@id="header"]/div/div[2]/h1')
            property_text = property_elem[0].text.strip() if property_elem else "N/A"
            
            duration_elem = dom.xpath('//*[@id="header"]/div/p')
            duration_text = duration_elem[0].text.strip() if duration_elem else "N/A"
            
            price_elem = dom.xpath('//*[@id="price"]/div/div/span/span[3]')
            price_text = price_elem[0].text.strip() if price_elem else "N/A"
            
            writer.writerow([url, property_text, duration_text, price_text])
            print(f"成功处理:{url}")
            
        except Exception as e:
            # 记录错误信息与对应URL
            error_msg = f"处理{url}时出错:{str(e)}"
            print(error_msg)
            # 可选:将错误写入日志文件
            # with open('error_log.txt', 'a', encoding='utf-8') as logfile:
            #     logfile.write(error_msg + '\n')

额外建议

  • 爬取大量页面时,加入随机延迟(比如time.sleep(random.uniform(1,3))),避免触发目标网站的反爬机制
  • 对于频繁变化的页面,建议定期检查XPath表达式是否仍然有效

内容的提问来源于stack exchange,提问作者Johnny FlimFlam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 22:10:28