如何使用BeautifulSoup抓取网页表格并转换为pandas DataFrame?
实现代码及步骤
首先补全依赖导入,请求时添加UA头避免被站点反爬拦截:
import requests from bs4 import BeautifulSoup import pandas as pd from bs4 import Comment url = 'https://www.basketball-reference.com/leagues/NBA_2021.html' # 添加UA头模拟浏览器请求,避免403错误 headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'} page = requests.get(url, headers=headers) soup = BeautifulSoup(page.content, 'html.parser')
定位Per Game Stats表格
该站点部分表格会放在注释内容中,优先直接查找,找不到再从注释中提取:
table = soup.find('table', id='per_game_stats') # 直接查找为空时,从注释中提取表格 if not table: comments = soup.find_all(string=lambda text: isinstance(text, Comment)) for comment in comments: if 'per_game_stats' in comment: comment_soup = BeautifulSoup(comment, 'html.parser') table = comment_soup.find('table', id='per_game_stats') break
提取表头和行数据
# 提取表头 headers = [th.get_text() for th in table.find('thead').find_all('th')] # 提取表格内容,跳过中间的分隔行 rows = [] for tr in table.find('tbody').find_all('tr'): # 跳过分类分隔行 if tr.get('class') and 'thead' in tr.get('class'): continue row = [tr.find('th').get_text()] + [td.get_text() for td in tr.find_all('td')] rows.append(row)
转换为pandas DataFrame
df = pd.DataFrame(rows, columns=headers) # 可选:将可转换的列转为数值类型 df = df.apply(pd.to_numeric, errors='ignore')
内容的提问来源于stack exchange,提问作者MKHN15
相关产品推荐
相关产品推荐

