You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬取Wunderground天气日历无结果求助

解决Wunderground日历爬取无结果的问题

问题原因

Wunderground的天气预报日历内容是JavaScript动态渲染的,你用requests.get()获取的只是页面的静态HTML骨架,动态加载的日历数据并没有包含在其中,所以BeautifulSoup找不到目标元素。

解决方案

方法1:使用Selenium模拟浏览器渲染

Selenium可以模拟真实浏览器加载页面,等待JavaScript执行完成后再获取完整的页面内容。

步骤:

  1. 安装Selenium:pip install selenium
  2. 下载对应浏览器的驱动(比如ChromeDriver,需与浏览器版本匹配)
  3. 修改代码如下:
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup

URL = "https://www.wunderground.com/calendar/us/ca/santa-barbara/KSBA"

# 初始化Chrome浏览器驱动
driver = webdriver.Chrome()
driver.get(URL)

# 等待目标元素加载完成,最多等待10秒
try:
    WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.CLASS_NAME, "calendar-days"))
    )
    # 获取渲染后的完整页面源码
    page_source = driver.page_source
    soup = BeautifulSoup(page_source, "html.parser")
    
    for ultag in soup.find_all('ul', {'class': 'calendar-days'}):
        for litag in ultag.find_all('li'):
            print(litag.text.strip())
finally:
    driver.quit()

方法2:直接调用网站API(更高效)

打开浏览器开发者工具(F12),切换到「网络」标签,刷新页面后,查找加载日历数据的XHR类型请求,找到对应的API地址后,直接用requests请求接口获取JSON格式数据,无需解析HTML。

示例(需自行验证真实API地址):

import requests

# 替换为实际找到的API地址
API_URL = "https://api.wunderground.com/api/[your-api-key]/calendar/q/CA/Santa_Barbara.json"
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
}

response = requests.get(API_URL, headers=headers)
if response.status_code == 200:
    data = response.json()
    # 提取并处理日历数据
    print(data)

额外注意事项

  • 设置合理的User-Agent请求头,避免被网站识别为爬虫
  • 部分网站有反爬机制,可能需要处理Cookie、请求频率限制等问题

内容的提问来源于stack exchange,提问作者Allen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 21:05:18