You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python网页爬取出现IndexError:列表索引越界问题求助

解决批量爬取URL文本时的IndexError问题

问题场景

批量遍历DataFrame中的URL爬取页面内容时,触发IndexError: list index out of range,报错位置在访问content[0]处。核心原因是部分页面不存在目标class元素,直接索引访问空列表导致报错。

报错信息

IndexError                                Traceback (most recent call last)
Input In [36], in <cell line: 3>()
      9 soup=BeautifulSoup(page.content,'html.parser')#parsing url text
     10 content=soup.findAll(attrs={'class':'td-post-content'})#extracting only text part
---&gt; 11 content=content[0].text.replace('\xa0',"  ").replace('\n',"  ")#replace end line symbol with space 
     12 title=soup.findAll(attrs={'class':'entry-title'})#extracting title of website
     13 title=title[16].text.replace('\n',"  ").replace('/','')

IndexError: list index out of range

问题根源

  1. soup.findAll()返回匹配元素的列表,若页面无对应class的元素,列表为空,直接访问[0]会触发索引越界
  2. 代码中title=title[16]同样存在风险:即使页面有entry-title元素,也未必有17个(索引从0开始),同样可能触发越界

修复方案

  • 用soup.find()替代soup.findAll()获取单个目标元素(找不到时返回None,避免空列表问题)
  • 添加判断逻辑,处理元素不存在的情况(如跳过当前URL、记录错误日志)
  • 避免硬编码索引(如title[16]),先判断列表长度再访问

修复后的代码

# 批量提取URL文本
import requests
from bs4 import BeautifulSoup
import pandas as pd
import numpy as np

url_id = 1
for i in range(len(df)):
    j = df.iloc[i].values
    headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/74.0.3729.169 Safari/537.36'}
    
    try:
        page = requests.get(j[0], headers=headers)
        page.raise_for_status()  # 捕获HTTP请求错误(如404、500)
        soup = BeautifulSoup(page.content, 'html.parser')
        
        # 获取内容:用find替代findAll,避免空列表
        content_elem = soup.find(attrs={'class': 'td-post-content'})
        if not content_elem:
            print(f"URL {j[0]} 未找到td-post-content元素,跳过")
            url_id += 1
            continue
        content = content_elem.text.replace('\xa0', "  ").replace('\n', "  ")
        
        # 获取标题:先判断元素数量是否足够,再访问指定索引
        title_elems = soup.findAll(attrs={'class': 'entry-title'})
        if len(title_elems) < 17:
            print(f"URL {j[0]} 的entry-title元素不足17个,跳过")
            url_id += 1
            continue
        title = title_elems[16].text.replace('\n', "  ").replace('/', '')
        
        # 合并内容并保存
        text = title + '.' + content
        df1 = pd.Series([text])
        filename = f"{url_id}.txt"
        # df1.to_csv(filename, line_terminator=',', index=False, header=False)
        # files.download(filename)
        print(f"已处理URL {url_id}: {j[0]}")
        url_id += 1
        
    except Exception as e:
        print(f"处理URL {j[0]} 时出错: {str(e)}")
        url_id += 1

关键改进点

  • 增加try-except块捕获请求和解析过程中的所有异常
  • 检查元素是否存在/数量是否足够,避免直接索引空列表或长度不足的列表
  • 添加日志输出,便于定位问题URL
  • 调用page.raise_for_status()捕获HTTP请求失败的情况

内容的提问来源于stack exchange,提问作者Binod Binod

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 01:20:38