求助解决Python爬虫中的IndexError: list index out of range错误
问题排查:爬虫脚本IndexError报错解决
报错原因
IndexError: list index out of range 说明你的代码中property、duration或price这三个变量里,至少有一个是空列表——也就是对应的XPath表达式在当前爬取的页面中没有找到目标元素。可能的诱因包括:
- 目标页面的HTML结构发生了变化,原XPath路径失效
- 请求返回的内容异常(比如404页面、反爬拦截页面)
- XPath写法过于依赖固定的DOM层级,容错性差
修复方案
1. 增加空值检查与错误记录
在尝试获取元素文本前,先判断列表是否为空,避免直接索引空列表;同时记录出错的URL,方便后续针对性排查。
2. 优化XPath表达式
尽量避免使用依赖固定层级的绝对路径,改用元素的特征(比如class属性、文本内容)来定位,提升鲁棒性。例如:
- 原
//*[@id="header"]/div/div[2]/h1可改为//h1[contains(@class, 'property-title')](假设标题有对应的class) - 原
//*[@id="price"]/div/div/span/span[3]可改为//span[contains(@class, 'price-value')](假设价格元素有标识class)
3. 增加请求异常处理
捕获请求过程中可能出现的错误(比如网络超时、连接失败),避免脚本直接崩溃。
4. 优化文件写入逻辑
不要在循环内反复打开/关闭输出文件,改为在循环外打开,提升运行效率。
修改后的完整代码
import requests from bs4 import BeautifulSoup from lxml import etree import csv # 提前打开输出文件,避免循环内重复IO操作 with open('1_colonia.csv', 'r', encoding='utf-8') as infile, \ open('2_colonia.csv', 'a', newline='', encoding='utf-8') as outfile: reader = csv.reader(infile, delimiter=';') writer = csv.writer(outfile, delimiter=';') # 可选:如果输出文件是新的,先写入表头 # writer.writerow(["URL", "Property", "Duration", "Price"]) next(reader) # 跳过表头 for row in reader: url = row[0] try: # 增加请求超时与异常捕获 page = requests.get(url, timeout=10) page.raise_for_status() # 检查请求是否成功(比如404、500会抛出异常) soup = BeautifulSoup(page.content, 'html.parser') dom = etree.HTML(str(soup)) # 用更鲁棒的XPath,同时处理空列表情况 property_elem = dom.xpath('//*[@id="header"]/div/div[2]/h1') property_text = property_elem[0].text.strip() if property_elem else "N/A" duration_elem = dom.xpath('//*[@id="header"]/div/p') duration_text = duration_elem[0].text.strip() if duration_elem else "N/A" price_elem = dom.xpath('//*[@id="price"]/div/div/span/span[3]') price_text = price_elem[0].text.strip() if price_elem else "N/A" writer.writerow([url, property_text, duration_text, price_text]) print(f"成功处理:{url}") except Exception as e: # 记录错误信息与对应URL error_msg = f"处理{url}时出错:{str(e)}" print(error_msg) # 可选:将错误写入日志文件 # with open('error_log.txt', 'a', encoding='utf-8') as logfile: # logfile.write(error_msg + '\n')
额外建议
- 爬取大量页面时,加入随机延迟(比如
time.sleep(random.uniform(1,3))),避免触发目标网站的反爬机制 - 对于频繁变化的页面,建议定期检查XPath表达式是否仍然有效
内容的提问来源于stack exchange,提问作者Johnny FlimFlam
相关产品推荐
相关产品推荐

