You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

从Tutiempo抓取太阳辐射数据遇迭代问题求助

问题修正方案

你的代码存在几个关键问题导致无法正常提取数据,以下是修复后的版本及问题说明:

核心问题点

  • 循环中误用列表变量而非当前迭代元素:你在循环里调用了solar_radiation_div(这是find_all返回的所有div的列表),而非循环变量solar_radiation_divs,导致每次循环都操作整个列表而非单个小时的div。
  • 未初始化data列表:直接调用data.append()前没有定义data = [],会触发NameError。
  • 缺少元素存在性检查:如果某个div里没有目标span,直接调用.text会报错,建议增加判断逻辑。

修复后的代码

import requests
from bs4 import BeautifulSoup

# URL of the website to scrape
url = "https://en.tutiempo.net/solar-radiation/lethbridge.html"

# Send a GET request to the URL
response = requests.get(url)

# Create a BeautifulSoup object to parse the HTML content
soup = BeautifulSoup(response.content, "html.parser")

# Find the div that contains the solar radiation data
solar_radiation_divs = soup.find_all("div", class_="horashidsow ocultas")

# Initialize data list to store results
data = []

# Extract the solar radiation data from each hourly sub-div
for hour_div in solar_radiation_divs:
    # Extract the hour from the sub-div's ID
    hour = hour_div["id"].split("-")[1]
    
    # Extract the solar radiation data with existence check
    ener_span = hour_div.find("span", class_="ener")
    if ener_span:
        solar_radiation = ener_span.text.strip()
        print(f"Hour {hour}: {solar_radiation} W/m²")
        data.append(solar_radiation)
    else:
        print(f"Hour {hour}: No solar radiation data found")
        data.append(None)

# Optional: Print collected data
print("Collected data:", data)

代码说明

  1. 变量重命名:将循环变量改为hour_div,更清晰表示当前处理的是单个小时的div元素。
  2. 初始化数据列表:在循环前定义data = [],避免未定义错误。
  3. 元素存在性检查:先用find获取span元素,确认存在后再提取文本,避免因元素缺失引发AttributeError。
  4. 文本清理:用.strip()去除文本前后空白字符,让数据更整洁。

内容的提问来源于stack exchange,提问作者user20275872

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 12:25:21