Python爬取Wunderground天气日历无结果求助
解决Wunderground日历爬取无结果的问题
问题原因
Wunderground的天气预报日历内容是JavaScript动态渲染的,你用requests.get()获取的只是页面的静态HTML骨架,动态加载的日历数据并没有包含在其中,所以BeautifulSoup找不到目标元素。
解决方案
方法1:使用Selenium模拟浏览器渲染
Selenium可以模拟真实浏览器加载页面,等待JavaScript执行完成后再获取完整的页面内容。
步骤:
- 安装Selenium:
pip install selenium - 下载对应浏览器的驱动(比如ChromeDriver,需与浏览器版本匹配)
- 修改代码如下:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup URL = "https://www.wunderground.com/calendar/us/ca/santa-barbara/KSBA" # 初始化Chrome浏览器驱动 driver = webdriver.Chrome() driver.get(URL) # 等待目标元素加载完成,最多等待10秒 try: WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CLASS_NAME, "calendar-days")) ) # 获取渲染后的完整页面源码 page_source = driver.page_source soup = BeautifulSoup(page_source, "html.parser") for ultag in soup.find_all('ul', {'class': 'calendar-days'}): for litag in ultag.find_all('li'): print(litag.text.strip()) finally: driver.quit()
方法2:直接调用网站API(更高效)
打开浏览器开发者工具(F12),切换到「网络」标签,刷新页面后,查找加载日历数据的XHR类型请求,找到对应的API地址后,直接用requests请求接口获取JSON格式数据,无需解析HTML。
示例(需自行验证真实API地址):
import requests # 替换为实际找到的API地址 API_URL = "https://api.wunderground.com/api/[your-api-key]/calendar/q/CA/Santa_Barbara.json" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } response = requests.get(API_URL, headers=headers) if response.status_code == 200: data = response.json() # 提取并处理日历数据 print(data)
额外注意事项
- 设置合理的
User-Agent请求头,避免被网站识别为爬虫 - 部分网站有反爬机制,可能需要处理Cookie、请求频率限制等问题
内容的提问来源于stack exchange,提问作者Allen
相关产品推荐
相关产品推荐

