使用BeautifulSoup无法读取网页表格数据的问题求助
解决Our World in Data表格爬取失败的问题
问题原因
该网站的表格内容是通过JavaScript动态渲染生成的,直接使用requests获取的原始HTML中并不包含表格数据,因此BeautifulSoup无法定位到目标表格元素。
解决方案1:直接调用官方API获取数据(推荐)
Our World in Data提供了公开的数据API,你可以直接请求接口获取结构化数据,无需解析HTML:
import pandas as pd import requests # 该数据集的API接口地址 api_url = "https://api.ourworldindata.org/v1/indicators/air-pollution-deaths-from-fossil-fuels.data.json" response = requests.get(api_url) data = response.json() # 将数据转换为DataFrame df = pd.DataFrame(data['data']) print(df.head())
说明:API返回的是JSON格式的原始数据,包含所有年份和国家的记录,你可以根据需求筛选或处理数据。
解决方案2:使用Selenium渲染动态页面
如果需要模拟浏览器行为获取渲染后的HTML,可以使用Selenium:
- 先安装Selenium和对应浏览器的驱动(如ChromeDriver)
- 运行以下代码:
from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.chrome.options import Options import pandas as pd # 配置Chrome无头模式 chrome_options = Options() chrome_options.add_argument("--headless=new") driver = webdriver.Chrome(options=chrome_options) url = "https://ourworldindata.org/grapher/pollution-deaths-from-fossil-fuels?tab=table" driver.get(url) # 等待页面加载完成(可根据实际情况调整等待时间) driver.implicitly_wait(10) # 获取渲染后的HTML soup = BeautifulSoup(driver.page_source, 'html.parser') driver.quit() # 提取表格数据 table = soup.select_one("table.data-table") df = pd.read_html(str(table))[0] print(df.head())
内容的提问来源于stack exchange,提问作者beridzeg45
相关产品推荐
相关产品推荐

