You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将网页抓取表格转为DataFrame?解决索引越界等报错问题

问题解决思路及代码修改

首先明确几个关键问题:

  • pd.read_html()返回的是DataFrame对象的列表,直接把这个列表传给pd.DataFrame()会生成3维结构,所以报错Must pass 2-d input。
  • 用dfs[1]索引越界,是因为当前dfs里只有1个元素(索引为0),说明默认情况下read_html只识别到了网页里的1个表格。

解决步骤

  1. 定位目标表格:打开目标网页,右键选择「检查」,找到你需要的第二个表格的<table>标签,查看它的id或class属性(比如目标表格可能有id="pitching"这类标识)。
  2. 精准抓取表格:给pd.read_html()添加attrs参数,指定目标表格的属性,直接抓取你要的表格,不用管列表索引。

修改后的代码示例

import requests
import pandas as pd

# 请求网页
response = requests.get('https://www.baseball-reference.com/players/p/penaje02.shtml')
# 替换成你找到的目标表格的id或class属性
target_table = pd.read_html(response.content, attrs={'id': 'pitching'})
# 因为指定了唯一属性,返回的列表只有1个元素,直接取索引0
df = target_table[0]
# 查看结果
print(df.head())

补充说明

如果目标表格没有明确的id或class,可以用BeautifulSoup先解析网页提取表格代码,再传给pd.read_html():

import requests
import pandas as pd
from bs4 import BeautifulSoup

response = requests.get('https://www.baseball-reference.com/players/p/penaje02.shtml')
soup = BeautifulSoup(response.text, 'html.parser')
# 找到页面中的第二个表格(索引从0开始,所以用find_all('table')[1])
target_table_html = soup.find_all('table')[1]
df = pd.read_html(str(target_table_html))[0]
print(df.head())

内容的提问来源于stack exchange,提问作者WilC

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 08:50:42