使用BeautifulSoup网页爬取时遭遇Index Error的解决方法
解决BeautifulSoup爬取数据时的IndexError问题
老哥,你的问题太典型了——爬取大量数据时突然中断触发IndexError,本质是你代码里直接用了[0]、[1]这种硬索引去取选择器返回的结果,但总有个别页面的结构和你预期的不一样,导致列表为空或者长度不够,直接取索引就炸了。
咱们一步步来搞定这个问题:
问题根源分析
看你代码里这些关键行:
car_make = car_soup.find("div",attrs={"class":"cldt-categorized-data cldt-data-section sc-pull-right"}).select("dl > dd:nth-of-type(1)")[0].text car_model = car_soup.find("div",attrs={"class":"cldt-categorized-data cldt-data-section sc-pull-right"}).select("dl > dd > a")[0].text car_year = car_soup.find("div",attrs={"class":"cldt-categorized-data cldt-data-section sc-pull-right"}).select("dl > dd > a")[1].text
一旦某个车辆详情页的结构和你预设的不一样(比如少了某个dd标签、a标签数量不足,甚至整个数据区域的class名微调),select()返回的列表长度就会小于你要取的索引,直接抛出IndexError,程序直接中断。
具体解决方案
1. 增加异常捕获,做容错处理
在处理单辆车详情的循环里加try-except块,遇到错误就跳过当前车辆,还能顺便记录错误信息方便后续排查。
2. 先检查元素存在性,再访问索引
别直接链式调用加索引,先把选择器结果存到变量里,判断长度够不够再取值,避免硬索引翻车。
3. 优化选择器,基于文本定位更可靠
别依赖nth-of-type或者固定位置的标签,而是先找到包含"Make"、"Model"等文本的dt标签,再取对应的dd内容——哪怕页面结构微调,只要文本标识不变,就不容易出错。
修改后的完整代码
import requests from bs4 import BeautifulSoup import time totalCar = 0 for pageNumber in range(3, 7): # 用f-string拼接URL更简洁 url = f"https://www.autoscout24.com/lst/bmw?sort=standard&desc=0&offer=U&ustate=N%2CU&size=20&page={pageNumber}&cy=D&mmm=47%7C%7C&mmm=9%7C%7C&atype=C&" r = requests.get(url) # 先检查页面请求是否成功 if r.status_code != 200: print(f"第{pageNumber}页请求失败,状态码:{r.status_code}") continue soup = BeautifulSoup(r.content,"lxml") car_details = soup.find_all("div",attrs={"class":"cl-list-element cl-list-element-gap"}) for detail in car_details: try: car_link = "https://www.autoscout24.com" + detail.a.get("href") # 加个延迟,避免请求太频繁被封 time.sleep(1) car_r = requests.get(car_link) if car_r.status_code != 200: print(f"车辆链接{car_link}请求失败") continue car_soup = BeautifulSoup(car_r.content,"lxml") data_section = car_soup.find("div",attrs={"class":"cldt-categorized-data cldt-data-section sc-pull-right"}) # 先判断数据区域是否存在 if not data_section: print(f"车辆{car_link}未找到核心数据区域") continue # 初始化变量 car_make = car_model = car_year = car_color = car_body = "" # 遍历所有dl标签对,根据dt文本匹配属性 dl_pairs = data_section.find_all("dl") for dl in dl_pairs: dt_text = dl.find("dt").text.strip() dd = dl.find("dd") if not dd: continue if dt_text == "Make": car_make = dd.text.strip() elif dt_text == "Model": car_model = dd.text.strip() elif dt_text == "Year": car_year = dd.text.strip() elif dt_text == "Color": car_color = dd.text.strip() elif dt_text == "Body": car_body = dd.text.strip() # 只输出有有效数据的条目 if car_make and car_model: print(f"Make:{car_make} Model:{car_model} Year:{car_year} Color:{car_color} Body:{car_body}") print("-"*20) totalCar += 1 except Exception as e: print(f"处理车辆时出错:{str(e)},链接:{car_link}") continue print(f"总共爬取到{totalCar}辆车")
额外小建议
- 一定要加请求延迟(比如代码里的
time.sleep(1)),避免短时间内请求太频繁被网站封禁IP - 可以把错误信息写到日志文件里,方便后续针对性排查异常页面
内容的提问来源于stack exchange,提问作者user14189147
相关产品推荐
相关产品推荐

