You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup无法读取网页表格数据的问题求助

解决Our World in Data表格爬取失败的问题

问题原因

该网站的表格内容是通过JavaScript动态渲染生成的,直接使用requests获取的原始HTML中并不包含表格数据,因此BeautifulSoup无法定位到目标表格元素。

解决方案1:直接调用官方API获取数据(推荐)

Our World in Data提供了公开的数据API,你可以直接请求接口获取结构化数据,无需解析HTML:

import pandas as pd
import requests

# 该数据集的API接口地址
api_url = "https://api.ourworldindata.org/v1/indicators/air-pollution-deaths-from-fossil-fuels.data.json"
response = requests.get(api_url)
data = response.json()

# 将数据转换为DataFrame
df = pd.DataFrame(data['data'])
print(df.head())

说明:API返回的是JSON格式的原始数据,包含所有年份和国家的记录,你可以根据需求筛选或处理数据。

解决方案2:使用Selenium渲染动态页面

如果需要模拟浏览器行为获取渲染后的HTML,可以使用Selenium:

  1. 先安装Selenium和对应浏览器的驱动(如ChromeDriver)
  2. 运行以下代码:
from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
import pandas as pd

# 配置Chrome无头模式
chrome_options = Options()
chrome_options.add_argument("--headless=new")
driver = webdriver.Chrome(options=chrome_options)

url = "https://ourworldindata.org/grapher/pollution-deaths-from-fossil-fuels?tab=table"
driver.get(url)

# 等待页面加载完成(可根据实际情况调整等待时间)
driver.implicitly_wait(10)

# 获取渲染后的HTML
soup = BeautifulSoup(driver.page_source, 'html.parser')
driver.quit()

# 提取表格数据
table = soup.select_one("table.data-table")
df = pd.read_html(str(table))[0]
print(df.head())

内容的提问来源于stack exchange,提问作者beridzeg45

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 16:36:11