如何从网页中抓取表格?Python爬取百老汇演出统计数据遇行数据填充失败问题
修复方案
你代码存在3处明显问题,调整后即可正常抓取数据:
- 变量名不匹配:解析后的soup对象命名为
springsteen_soup,查找表格时误用了未定义的soup;提取的第一个表格存为table1,遍历行时误用了未定义的table - 未导入pandas库,代码里直接使用
pd会报错 - pandas高版本已废弃
DataFrame.append()方法,且部分单元格取span.contents的逻辑容错性不足,容易触发异常
import requests import pandas as pd from bs4 import BeautifulSoup # 加请求头避免被网站反爬拦截 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36' } springsteen_url = 'https://www.ibdb.com/broadway-production/springsteen-on-broadway-515480#Statistics' springsteen_response = requests.get(springsteen_url, headers=headers) springsteen_soup = BeautifulSoup(springsteen_response.text, 'html.parser') # 提取第一个striped类表格 statistics = springsteen_soup.find_all("table", attrs={"class": "striped"}) table1 = statistics[0] # 定义数据表 df = pd.DataFrame(columns=['Week Ending', 'Gross', '% Gross Pot.', 'Attendance', '% Capacity']) rows_data = [] # 遍历表格行提取数据 for row in table1.tbody.find_all('tr'): columns = row.find_all('td') if columns: week = columns[0].text.strip() gross = columns[1].text.strip() # 优化取值逻辑,避免span不存在时报错,同时清理多余符号方便后续转数值 grosspot = columns[2].text.strip().replace('%', '') attendance = columns[3].text.strip().replace(',', '') capacity = columns[4].text.strip().replace('%', '') rows_data.append({ 'Week Ending': week, 'Gross': gross, '% Gross Pot.': grosspot, 'Attendance': attendance, '% Capacity': capacity }) # 批量写入DataFrame,效率更高 df = pd.concat([df, pd.DataFrame(rows_data)], ignore_index=True) print(df.head())
内容的提问来源于stack exchange,提问作者Galen04
相关产品推荐
相关产品推荐

