You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用BeautifulSoup抓取网页表格并转换为pandas DataFrame?

实现代码及步骤

首先补全依赖导入,请求时添加UA头避免被站点反爬拦截:

import requests
from bs4 import BeautifulSoup
import pandas as pd
from bs4 import Comment

url = 'https://www.basketball-reference.com/leagues/NBA_2021.html'
# 添加UA头模拟浏览器请求,避免403错误
headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'}
page = requests.get(url, headers=headers)
soup = BeautifulSoup(page.content, 'html.parser')

定位Per Game Stats表格

该站点部分表格会放在注释内容中,优先直接查找,找不到再从注释中提取:

table = soup.find('table', id='per_game_stats')

# 直接查找为空时,从注释中提取表格
if not table:
    comments = soup.find_all(string=lambda text: isinstance(text, Comment))
    for comment in comments:
        if 'per_game_stats' in comment:
            comment_soup = BeautifulSoup(comment, 'html.parser')
            table = comment_soup.find('table', id='per_game_stats')
            break

提取表头和行数据

# 提取表头
headers = [th.get_text() for th in table.find('thead').find_all('th')]

# 提取表格内容,跳过中间的分隔行
rows = []
for tr in table.find('tbody').find_all('tr'):
    # 跳过分类分隔行
    if tr.get('class') and 'thead' in tr.get('class'):
        continue
    row = [tr.find('th').get_text()] + [td.get_text() for td in tr.find_all('td')]
    rows.append(row)

转换为pandas DataFrame

df = pd.DataFrame(rows, columns=headers)
# 可选:将可转换的列转为数值类型
df = df.apply(pd.to_numeric, errors='ignore')

内容的提问来源于stack exchange,提问作者MKHN15

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 09:00:00