如何使用BeautifulSoup提取谷歌搜索结果SVG图表内的小时天气数据?
问题根因
谷歌搜索的逐小时预报SVG内部元素由页面JavaScript动态渲染生成,requests仅能获取未执行JS的原始静态HTML,因此返回的SVG节点为空。你在浏览器控制台看到的完整DOM是浏览器完成JS渲染后的结果,和初始请求的静态HTML内容不一致。
解决方案
这里提供两种可行路径,优先推荐第一种轻量方案:
方案1:直接提取静态HTML内置的结构化逐小时数据(无需解析SVG)
谷歌天气页面的原始HTML已经内置了全量逐小时预报的结构化数据,无需解析SVG即可直接提取,调整代码如下:
from bs4 import BeautifulSoup as bs import requests def get_weather_data(region): USER_AGENT = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/95.0.4638.54 Safari/537.36" LANGUAGE = "en-US,en;q=0.5" URL = f"https://www.google.com/search?lr=lang_en&q=weather+in+{region.strip().lower().replace(' ', '+')}" s = requests.Session() s.headers['User-Agent'] = USER_AGENT s.headers['Accept-Language'] = LANGUAGE s.headers['Content-Language'] = LANGUAGE html = s.get(URL) soup = bs(html.text, "html.parser") # 提取逐小时预报数据 hourly_forecast = [] # 逐小时条目容器 hourly_items = soup.find_all("div", class_="wob_hf") for item in hourly_items: hour = item.find("div", class_="wob_h").get_text(strip=True) temp = item.find("span", class_="wob_t").get_text(strip=True) weather_desc = item.find("img")["alt"] precipitation = item.find("span", class_="wob_pp").get_text(strip=True) humidity = item.find("span", class_="wob_hm").get_text(strip=True) wind = item.find("span", class_="wob_ws").get_text(strip=True) hourly_forecast.append({ "小时": hour, "温度": temp, "天气描述": weather_desc, "降水概率": precipitation, "湿度": humidity, "风速": wind }) return hourly_forecast print(get_weather_data("London"))
方案2:使用支持JS渲染的工具获取完整SVG内容
如果确实需要提取SVG内部的绘图数据,可以使用支持浏览器内核渲染的工具执行页面JS后再提取DOM,示例使用Selenium:
from bs4 import BeautifulSoup as bs from selenium import webdriver from selenium.webdriver.chrome.options import Options def get_weather_svg(region): chrome_options = Options() chrome_options.add_argument("--headless") # 无头模式不弹出浏览器 chrome_options.add_argument(f"user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/95.0.4638.54 Safari/537.36") chrome_options.add_argument("accept-language=en-US,en;q=0.5") driver = webdriver.Chrome(options=chrome_options) URL = f"https://www.google.com/search?lr=lang_en&q=weather+in+{region.strip().lower().replace(' ', '+')}" driver.get(URL) # 等待JS渲染完成 driver.implicitly_wait(3) soup = bs(driver.page_source, "html.parser") hourly_svg = soup.find("svg", attrs={'id':'wob_gsvg'}) print(hourly_svg.prettify()) driver.quit() get_weather_svg("London")
注意使用Selenium需要提前安装对应版本的ChromeDriver和依赖库。
内容的提问来源于stack exchange,提问作者Curious Learner
相关产品推荐
相关产品推荐

