You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

爬取IPL官网2020赛季数据触发IndexError的解决求助

错误原因

  • 核心问题是BeautifulSoup的findAll方法参数使用错误:你把标签名和class属性合并写在了name参数里,name参数仅用于指定HTML标签名称,不能直接拼接属性,所以该查询返回空列表,取索引[0]时触发IndexError。
  • 次要注意点:findAll是BeautifulSoup3的旧写法,BS4更推荐使用find_all方法,兼容性更好。

修复后的代码

import json
import pandas as pd
from bs4 import BeautifulSoup
from urllib.request import urlopen

scrape_url="https://www.iplt20.com/stats/2020/most-runs"
page_connect = urlopen(scrape_url)

page_html=BeautifulSoup(page_connect, 'html.parser')

# 正确写法:单独指定标签名和class属性
target_divs = page_html.find_all(name="div", class_="js-table")
# 先判断是否找到元素再操作,避免索引报错
if target_divs:
    json_raw_string = target_divs[0].string
    print(json_raw_string)
else:
    print("未找到对应属性的div元素")

额外注意事项

  • 部分动态渲染的内容如果通过静态请求无法获取,你可以进一步检查页面返回的HTML结构,确认js-table类的div是否存在于静态源码中。
  • 爬取前建议先检查目标网站的robots协议,确认爬取行为符合站点规则。

内容的提问来源于stack exchange,提问作者Raksha

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.25 05:45:04