You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从网页中抓取表格?Python爬取百老汇演出统计数据遇行数据填充失败问题

修复方案

你代码存在3处明显问题,调整后即可正常抓取数据:

  • 变量名不匹配:解析后的soup对象命名为springsteen_soup,查找表格时误用了未定义的soup;提取的第一个表格存为table1,遍历行时误用了未定义的table
  • 未导入pandas库,代码里直接使用pd会报错
  • pandas高版本已废弃DataFrame.append()方法,且部分单元格取span.contents的逻辑容错性不足,容易触发异常
import requests
import pandas as pd
from bs4 import BeautifulSoup

# 加请求头避免被网站反爬拦截
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
}
springsteen_url = 'https://www.ibdb.com/broadway-production/springsteen-on-broadway-515480#Statistics'
springsteen_response = requests.get(springsteen_url, headers=headers)
springsteen_soup = BeautifulSoup(springsteen_response.text, 'html.parser')

# 提取第一个striped类表格
statistics = springsteen_soup.find_all("table", attrs={"class": "striped"})
table1 = statistics[0]

# 定义数据表
df = pd.DataFrame(columns=['Week Ending', 'Gross', '% Gross Pot.', 'Attendance', '% Capacity'])
rows_data = []

# 遍历表格行提取数据
for row in table1.tbody.find_all('tr'):
    columns = row.find_all('td')
    if columns:
        week = columns[0].text.strip()
        gross = columns[1].text.strip()
        # 优化取值逻辑,避免span不存在时报错,同时清理多余符号方便后续转数值
        grosspot = columns[2].text.strip().replace('%', '')
        attendance = columns[3].text.strip().replace(',', '')
        capacity = columns[4].text.strip().replace('%', '')
        rows_data.append({
            'Week Ending': week,
            'Gross': gross,
            '% Gross Pot.': grosspot,
            'Attendance': attendance,
            '% Capacity': capacity
        })

# 批量写入DataFrame,效率更高
df = pd.concat([df, pd.DataFrame(rows_data)], ignore_index=True)
print(df.head())

内容的提问来源于stack exchange,提问作者Galen04

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.27 03:15:05