You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

单站点可爬取wunderground数据存CSV,4站点批量爬取至桌面目录失败求助

批量爬取Wunderground站点数据解决方案

嘿,这事儿其实不难,核心就是把你现有的单站点爬取逻辑封装成可复用的函数,然后通过循环遍历你的站点列表和月份列表来批量执行。我给你梳理一套完整的Python实现思路,你可以直接套用到你的代码里:

步骤1:读取站点配置文件

首先用pandas或者原生Python读取你的stations.csv,把每个站点的信息存成一个列表方便遍历:

import pandas as pd
import os

# 读取stations.csv(注意如果是分号分隔要加sep=";")
stations_df = pd.read_csv("stations.csv", header=None, names=["station_id", "lat", "lon"], sep=";")
stations = stations_df.to_dict("records")  # 转成字典列表,每个元素是{"station_id": "...", "lat": ..., "lon": ...}

# 定义桌面保存路径(跨系统通用)
desktop_path = os.path.expanduser("~/Desktop")

步骤2:封装爬取与清洗函数

把你现在单站点的爬取逻辑抽成一个函数,接收station_id和year_month(比如"2023-01")作为参数,返回清洗后的DataFrame:

import requests
from bs4 import BeautifulSoup
import pandas as pd

def scrape_station_data(station_id, year_month):
    # 构造Wunderground月度历史数据URL
    year, month = year_month.split("-")
    url = f"https://www.wunderground.com/dashboard/pws/{station_id}/history/monthly/{year}/{month}/daily"
    
    # 模拟浏览器请求,避免被拦截
    headers = {
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
    }
    response = requests.get(url, headers=headers)
    response.raise_for_status()  # 请求失败时抛出异常
    
    # 这里替换成你实际的页面解析、数据清洗逻辑
    soup = BeautifulSoup(response.text, "html.parser")
    # ... 提取表格数据、清洗脏数据、整理成DataFrame ...
    cleaned_data = pd.DataFrame(...)  # 你的清洗后数据
    
    return cleaned_data

步骤3:批量遍历执行

嵌套循环遍历每个站点和12个月份,调用爬取函数并保存CSV:

import time

# 定义要爬取的年份,可根据需求修改
target_year = "2023"
# 生成12个月份的格式字符串:["2023-01", "2023-02", ..., "2023-12"]
months = [f"{target_year}-{str(m).zfill(2)}" for m in range(1, 13)]

for station in stations:
    station_id = station["station_id"]
    for month in months:
        try:
            print(f"正在爬取 {station_id} - {month} 的数据...")
            data = scrape_station_data(station_id, month)
            
            # 构造保存文件名,比如"KCASANFR131_2023-01.csv"
            filename = f"{station_id}_{month}.csv"
            save_path = os.path.join(desktop_path, filename)
            
            # 保存CSV(跳过空数据,避免生成空文件)
            if not data.empty:
                data.to_csv(save_path, index=False)
                print(f"已保存: {save_path}")
            else:
                print(f"{station_id} - {month} 无有效数据,跳过保存")
                
            # 加入延迟,避免触发反爬机制
            time.sleep(2)
            
        except Exception as e:
            print(f"爬取 {station_id} - {month} 失败: {str(e)}")
            # 记录错误日志到桌面,方便后续排查
            with open(os.path.join(desktop_path, "scrape_error_log.txt"), "a") as f:
                f.write(f"{station_id} - {month}: {str(e)}\n")

额外实用提示

  • 反爬优化:如果遇到IP被封,可以考虑加入随机延迟(time.sleep(random.uniform(1,3))),或者使用代理IP池
  • 数据验证:在保存前可以检查DataFrame的列数、数据量,确保数据符合预期
  • 断点续爬:可以维护一个已完成的任务列表,每次运行前先跳过已经成功保存的文件,避免重复爬取

这样一套逻辑跑下来,就能自动生成4个站点×12个月=48个CSV文件到你的桌面啦!

内容的提问来源于stack exchange,提问作者Ma_

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 10:10:05