You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup4提取深层嵌套表格中高低温天气数据求助

解决Wunderground历史天气数据提取问题(BeautifulSoup4 + 动态渲染)

问题根源

你遇到的核心问题是目标页面用Angular框架动态渲染内容,BeautifulSoup只能爬取页面加载时的静态HTML,而lib-city-history-observation内部的表格数据是浏览器加载后通过JS动态生成的,所以静态抓取拿不到后续节点,找ng-star-inserted类的元素也会返回空结果。

解决方案:用动态渲染工具获取完整页面

推荐用Selenium(入门友好)模拟浏览器加载页面,等JS渲染完成后再抓取HTML,之后用BeautifulSoup解析。下面是详细步骤:

1. 安装依赖

先安装需要的Python包:

pip install selenium beautifulsoup4 webdriver-manager

webdriver-manager会自动管理浏览器驱动,不用手动下载配置。

2. 单页面数据提取代码示例

from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from webdriver_manager.chrome import ChromeDriverManager
from bs4 import BeautifulSoup
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By

# 初始化Chrome浏览器
driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()))

# 目标页面URL
url = "https://www.wunderground.com/history/monthly/us/ma/east-boston/KBOS/date/1982-1"
driver.get(url)

# 等待页面渲染完成(显式等待比固定sleep更可靠)
wait = WebDriverWait(driver, 10)
wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, "lib-city-history-observation .observation-table")))

# 获取渲染后的完整HTML
page_source = driver.page_source
driver.quit()  # 关闭浏览器

# 用BeautifulSoup解析HTML
soup = BeautifulSoup(page_source, 'html.parser')

# 定位数据节点
observation_node = soup.find('lib-city-history-observation')
if observation_node:
    table_container = observation_node.find('div', class_='observation-table ng-star-inserted')
    if table_container:
        # 找到嵌套的温度表格
        temp_table = table_container.find('table').find('td').find('table')
        if temp_table:
            # 提取第1组高低温数据(对应你给的CSS选择器位置)
            first_temp_row = temp_table.find_all('tr')[1]
            high_temp = first_temp_row.find_all('td')[0].text.strip()
            low_temp = first_temp_row.find_all('td')[1].text.strip()
            print(f"1982年1月首个高温:{high_temp}")
            print(f"1982年1月首个低温:{low_temp}")

3. 批量处理思路

  • 把URL做成模板:https://www.wunderground.com/history/monthly/us/ma/east-boston/KBOS/date/{year}-{month}
  • 循环遍历需要的年份和月份,生成对应URL
  • 将数据提取逻辑封装成函数,每个URL调用一次,把结果保存到CSV或Excel文件

备选高效方案:直接调用API接口

打开浏览器开发者工具(F12),切换到「Network」标签页刷新页面,查找包含history或observation的XHR请求,大概率能找到直接返回结构化JSON数据的API接口。用requests直接请求API,不用渲染页面,速度更快。

常见问题排查

  • 若遇到登录弹窗:Wunderground部分数据需要登录,可通过Selenium模拟登录,或抓包获取登录后的Cookie,在API请求中带上。
  • 动态类名失效:ng-star-inserted是Angular生成的动态类,可能随版本变化,尽量用表格结构、表头文字等稳定特征定位元素,不要依赖动态类名。

内容的提问来源于stack exchange,提问作者SecHumanist

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 19:42:54