You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python BeautifulSoup爬取天气数据出现数据不符问题求助

问题原因分析
  • 元素定位错误:你当前代码中选择的forecast-briefly__name元素,对应的是页面中的单日最低气温(或非实时的预报气温),并非网页显示的当前实时气温,这是数据不符的核心原因。
  • 静态请求局限性:Yandex部分天气数据可能通过JavaScript动态渲染,requests.get仅能获取页面初始的静态HTML,无法捕获JS加载后的实时数据(不过此案例中优先是元素定位问题)。
解决方法

方法1:修正元素定位(静态请求可解决)

先打开目标网页,通过浏览器开发者工具查看当前实时气温对应的DOM结构,替换代码中的选择器。以下是修正后的示例代码:

import requests
from bs4 import BeautifulSoup

link = "https://yandex.com.tr/hava/duzce?via=hnav"
response = requests.get(link)
soup = BeautifulSoup(response.content, "html.parser")

# 定位当前实时气温的正确元素(以实际DOM结构为准,示例为temp__value)
current_temp = soup.find("span", class_="temp__value")
if current_temp:
    print(current_temp.text)
else:
    print("未找到目标气温元素,请检查DOM结构")

方法2:模拟浏览器加载动态数据

若目标数据完全由JavaScript动态生成,静态请求无法获取,则使用Selenium模拟浏览器行为:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

link = "https://yandex.com.tr/hava/duzce?via=hnav"
# 初始化浏览器驱动(需提前安装对应浏览器的驱动并配置环境变量)
driver = webdriver.Chrome()
driver.get(link)

try:
    # 等待目标元素加载完成,超时时间10秒
    current_temp = WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.CLASS_NAME, "temp__value"))
    )
    print(current_temp.text)
finally:
    driver.quit()

注意事项

  • 网页DOM结构可能随网站更新变化,需定期验证目标元素的选择器是否有效。
  • 使用Selenium时,需注意网站的反爬机制,避免触发验证限制。

内容的提问来源于stack exchange,提问作者İsmail Eren

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.07 17:15:00