多气象站小时温度DataFrame的多时间尺度均值循环实现方法咨询
Great question—repeating the same resampling code for every station is tedious, so let's fix that with a clean, reusable workflow. Here's how to do it efficiently:
First, Prep Your Data
Make sure every station's DataFrame has a datetime index (or a column you can convert to one). This is non-negotiable for pandas resampling. For example:
# For a single station's DataFrame import pandas as pd station_df['timestamp'] = pd.to_datetime(station_df['timestamp']) station_df.set_index('timestamp', inplace=True)
Do this for all your station datasets first—consistency here will make the rest smooth.
Step 1: Write a Reusable Aggregation Function
Instead of copying resampling code for each station, define one function that handles all the temporal averages. This is where you centralize your logic:
def calculate_temporal_means(df, temp_column='temperature'): # Compute all resampled averages in one go daily = df[temp_column].resample('D').mean().rename('daily_avg') weekly = df[temp_column].resample('W').mean().rename('weekly_avg') monthly = df[temp_column].resample('M').mean().rename('monthly_avg') annual = df[temp_column].resample('Y').mean().rename('annual_avg') # Combine into a single DataFrame for easy access return pd.concat([daily, weekly, monthly, annual], axis=1)
You can tweak the resampling codes (like 'W-MON' for weeks starting on Monday) or add more intervals right here if needed.
Step 2: Loop Through All Stations
Store your station DataFrames in a dictionary (keys = station names, values = DataFrames) for easy iteration. Then loop through and apply the function:
# Example: Dictionary holding all station data station_datasets = { 'Downtown': downtown_df, 'Uptown': uptown_df, 'Suburb': suburb_df } # Dictionary to store results for each station station_averages = {} # Loop and compute averages for station_name, df in station_datasets.items(): print(f"Processing {station_name}...") station_averages[station_name] = calculate_temporal_means(df) # Access results for a specific station print(station_averages['Downtown']['daily_avg'].head())
Optional: Combine Results for Cross-Station Comparison
If you want to compare averages across stations, merge all results into a multi-index DataFrame:
combined_results = pd.concat(station_averages.values(), keys=station_averages.keys(), axis=1) # Example: View daily averages for all stations print(combined_results.xs('daily_avg', level=1, axis=1).head())
Quick Tips
- Handle Missing Data: Add
dropna=Falsetoresample()if you want to keep periods with no data, or usedf.fillna(method='ffill')before resampling to fill gaps (use cautiously!). - Efficiency: Pandas resampling is optimized, so even with large datasets, this loop will run quickly—you're just applying pre-built logic to each dataset without redundant code.
- Scalability: Add new stations to the
station_datasetsdictionary, and the loop will handle them automatically without any extra code.
内容的提问来源于stack exchange,提问作者Antonio Bonilla

