You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何修复网页爬虫中的IndexError: list index out of range错误

嘿,这个索引越界的问题我爬取表格数据时也踩过坑!咱们来一步步拆解问题并修复它:

问题根源

你现在用page_soup.findAll('tr')获取了页面所有的<tr>元素,但页面里肯定存在一些非目标数据行(比如表头行、分页行、空行或者结构不同的行),这些行的<td>数量远少于你预期的7个。当你尝试用索引访问这些行的td时,自然会触发IndexError——因为列表里根本没有这么多元素。

另外你用del containers[8]硬删除第8个元素的做法非常不稳定,一旦页面结构微调,这个删除操作就会出错或者删错内容。

修复方案

我们需要做两件关键的事:

  • 精准筛选目标行:只抓取带有featured类的<tr>(你的示例行有featured even,对应奇数行应该是featured odd),过滤掉无关的行。
  • 增加安全检查:在访问td索引之前,先判断当前行的td数量是否符合要求,避免越界。

修改后的完整代码

import csv
from bs4 import BeautifulSoup

# 假设page_html是你获取到的页面内容
page_soup = BeautifulSoup(page_html, "html.parser")
# 只抓取带有featured类的tr,精准定位目标数据行
containers = page_soup.findAll('tr', class_=['featured even', 'featured odd'])

company_names = []
booth_numbers = []
categories = []
countries = []

print("generating csv")
with open('CompanyList.csv','w', newline='') as f:  # 加上newline=''避免csv出现空行
    csv_out = csv.writer(f)
    csv_out.writerow(["company_name", "booth_number", "category", "country"])
    
    for container in containers:
        cols = container.findAll("td")
        # 先检查td数量是否足够(我们需要用到索引0-5,所以至少要有6个td)
        if len(cols) >= 6:
            try:
                company_name = cols[1].find("a").text.strip()
                booth_number = cols[2].text.strip()
                category = cols[3].text.strip()
                country = cols[5].text.strip()
                
                company_names.append(company_name)
                booth_numbers.append(booth_number)
                categories.append(category)
                countries.append(country)
                
                csv_out.writerow([company_name, booth_number, category, country])
            except AttributeError:
                # 处理可能的子元素不存在的情况(比如某个a标签缺失)
                print(f"跳过格式异常的行: {container}")
                continue
        else:
            print(f"跳过td数量不足的行: {container}")
            continue

print('Done Writing to File')

额外说明

  • 我给代码加上了newline=''参数,这是Python写入csv文件的最佳实践,可以避免Windows系统下出现多余的空行。
  • 增加了try-except块来处理可能的AttributeError(比如某个公司名称的<a>标签缺失),让代码更健壮。
  • 所有文本都加上了.strip(),可以去除多余的空格、换行符,让csv数据更干净。

内容的提问来源于stack exchange,提问作者PythonFisher

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 07:47:57