从Tutiempo抓取太阳辐射数据遇迭代问题求助
问题修正方案
你的代码存在几个关键问题导致无法正常提取数据,以下是修复后的版本及问题说明:
核心问题点
- 循环中误用列表变量而非当前迭代元素:你在循环里调用了
solar_radiation_div(这是find_all返回的所有div的列表),而非循环变量solar_radiation_divs,导致每次循环都操作整个列表而非单个小时的div。 - 未初始化
data列表:直接调用data.append()前没有定义data = [],会触发NameError。 - 缺少元素存在性检查:如果某个div里没有目标span,直接调用
.text会报错,建议增加判断逻辑。
修复后的代码
import requests from bs4 import BeautifulSoup # URL of the website to scrape url = "https://en.tutiempo.net/solar-radiation/lethbridge.html" # Send a GET request to the URL response = requests.get(url) # Create a BeautifulSoup object to parse the HTML content soup = BeautifulSoup(response.content, "html.parser") # Find the div that contains the solar radiation data solar_radiation_divs = soup.find_all("div", class_="horashidsow ocultas") # Initialize data list to store results data = [] # Extract the solar radiation data from each hourly sub-div for hour_div in solar_radiation_divs: # Extract the hour from the sub-div's ID hour = hour_div["id"].split("-")[1] # Extract the solar radiation data with existence check ener_span = hour_div.find("span", class_="ener") if ener_span: solar_radiation = ener_span.text.strip() print(f"Hour {hour}: {solar_radiation} W/m²") data.append(solar_radiation) else: print(f"Hour {hour}: No solar radiation data found") data.append(None) # Optional: Print collected data print("Collected data:", data)
代码说明
- 变量重命名:将循环变量改为
hour_div,更清晰表示当前处理的是单个小时的div元素。 - 初始化数据列表:在循环前定义
data = [],避免未定义错误。 - 元素存在性检查:先用
find获取span元素,确认存在后再提取文本,避免因元素缺失引发AttributeError。 - 文本清理:用
.strip()去除文本前后空白字符,让数据更整洁。
内容的提问来源于stack exchange,提问作者user20275872
相关产品推荐
相关产品推荐

