You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python和Beautiful Soup抓取整数值?爬取温度失败求助

问题分析

你遇到的核心问题是:目标网站的温度数据是通过JavaScript动态渲染的。requests.get()只能获取页面的初始HTML源码,此时<span class="temp">标签还未被填充数据,所以BeautifulSoup无法提取到温度值。

解决方案

方案一:用Selenium模拟浏览器渲染页面

Selenium可以模拟真实浏览器加载页面,等待JavaScript执行完成后再获取完整的页面内容,这样就能拿到填充后的温度数据。

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup

url = 'https://www.theweathernetwork.com/us/weather/new-york/new-york'

# 初始化Chrome浏览器(需提前安装对应版本的ChromeDriver)
driver = webdriver.Chrome()
driver.get(url)

# 等待温度元素加载完成,最长等待10秒
wait = WebDriverWait(driver, 10)
wait.until(EC.presence_of_element_located((By.CLASS_NAME, 'temp')))

# 获取加载后的完整页面源码
page_source = driver.page_source
driver.quit()

# 解析并提取温度
soup = BeautifulSoup(page_source, 'lxml')
temp = soup.find('span', class_='temp').text
print(temp)

注意:需要先安装Selenium库(pip install selenium),并下载与浏览器版本匹配的驱动(如ChromeDriver)。

方案二:直接调用网站的API接口

通过浏览器开发者工具(F12→Network面板)抓包,可以找到网站获取温度数据的API接口,直接请求API能更高效地拿到数据,无需模拟浏览器。

示例代码(需替换为实际抓包到的API地址):

import requests

# 替换为实际抓到的温度API接口
api_url = "https://api.theweathernetwork.com/data/v1/weather?location=USNY0996"
response = requests.get(api_url)
data = response.json()

# 根据API返回的JSON结构提取温度(字段需以实际返回为准)
temp = data['observation']['temperature']['value']
print(f"{temp}°F")

注意:API接口可能会随网站更新变化,需要自行通过抓包工具确认最新的接口地址和参数。

内容的提问来源于stack exchange,提问作者Video

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 03:35:20