如何遍历气象数据目录文件,计算各站点月总降水量并输出CSV
Got it, let's build a Python script that handles this precipitation data aggregation task perfectly. I'll break it down step by step so you can follow along and adjust it to your exact CSV structure:
Python Script to Calculate Monthly Precipitation Totals for 4 Stations
First, we'll use pandas for this task—it's ideal for handling CSV data and aggregating totals. If you don't have it installed yet, run this in your terminal:
pip install pandas
The Script
import pandas as pd from pathlib import Path # Point to your database directory on the Desktop data_directory = Path.home() / "Desktop" / "database" # Empty list to collect all processed daily records daily_records = [] # Loop through every CSV file in the directory for csv_file in data_directory.glob("*.csv"): print(f"Processing: {csv_file.name}") # Read the CSV, extracting only the columns we need # --- IMPORTANT ADJUSTMENTS HERE --- # Assumptions about your CSV structure (0-based indices): # - Index 0: Station ID/Name (e.g., "Station_A", "Station_1") # - Index 1: Date (format like YYYY-MM-DD or YYYY/MM/DD) # - Index 16: Precipitation value (your target column) daily_data = pd.read_csv( csv_file, usecols=[0, 1, 16], header=None, # Delete this line if your CSV has a header row names=["Station", "Date", "Precipitation"] # Name the columns for clarity ) # Convert date column to datetime to extract month daily_data["Date"] = pd.to_datetime(daily_data["Date"], errors="coerce") # Remove rows with invalid dates (if any) daily_data = daily_data.dropna(subset=["Date"]) # Extract month number (1 = January, 12 = December) daily_data["Month"] = daily_data["Date"].dt.month # Add this file's data to our master list daily_records.append(daily_data) # Combine all daily data into one DataFrame all_data = pd.concat(daily_records, ignore_index=True) # Calculate total precipitation per station per month monthly_totals = all_data.groupby(["Station", "Month"], as_index=False)["Precipitation"].sum() # Rename the sum column for readability monthly_totals.rename(columns={"Precipitation": "Total_Precipitation"}, inplace=True) # Sort results by station and month to get clean 48-row output (4 stations × 12 months) monthly_totals = monthly_totals.sort_values(by=["Station", "Month"]).reset_index(drop=True) # Save the final result to a CSV on your Desktop output_file = Path.home() / "Desktop" / "monthly_precipitation_summary.csv" monthly_totals.to_csv(output_file, index=False) print(f"Done! Your 48-row summary is saved to: {output_file}")
Key Adjustments for Your Data
- CSV Header Row: If your CSV files have a header row (e.g., first line says "Station,Date,Temperature,..."), delete the
header=Noneline and updateusecolsto use column names instead of indices (e.g.,usecols=["Station", "ObservationDate", "DailyPrecip"]). - Non-Standard Date Format: If your dates are in a weird format (like
DD/MM/YYYY), add theformatparameter topd.to_datetime:pd.to_datetime(daily_data["Date"], format="%d/%m/%Y", errors="coerce"). - Station in Filename: If your CSV files don't have a station column but the filename includes the station ID (e.g., "station_2_20230515.csv"), extract it from the filename like this:
# Add this right after opening the CSV file station_id = csv_file.name.split("_")[1] # Adjust split logic to match your filename pattern daily_data["Station"] = station_id
This script will process all your CSV files, aggregate the monthly totals for each of the 4 stations, and output a clean 48-row CSV with the results.
内容的提问来源于stack exchange,提问作者Hun_
相关产品推荐
相关产品推荐

