Pandas高效遍历多DataFrame:多年份气温数据集处理求助
Hey Sergio, great to hear you're diving into Pandas with Jupyter for your temperature datasets! Let's work through the common pain points you've mentioned—column names with spaces, split date columns, and handling 10 separate files efficiently. Here's a step-by-step breakdown:
Those spaces in column names are going to be a nuisance down the line (trust me, typing df['Temperature Value'] every time gets old fast). Let's fix them right away when loading your data:
import pandas as pd # Load a single dataset as an example df = pd.read_csv("your_temp_data.csv") # Replace spaces with underscores and convert to lowercase (adjust to your preference) df.columns = df.columns.str.replace(" ", "_").str.lower() # If you know specific column names you want to rename, you can do this instead: # df.rename(columns={"Original Column Name": "clean_column_name"}, inplace=True)
This will turn something like "Daily Temperature" into daily_temperature, making it way easier to reference columns in your code.
Since your date data is split across separate columns (like agno for year, mes for month, I assume there's a dia column for day too), let's combine them into a proper datetime column—this is essential for time-series analysis:
# First, make sure your year/month/day columns are integer types (fix if needed) df['agno'] = df['agno'].astype(int) df['mes'] = df['mes'].astype(int) df['dia'] = df['dia'].astype(int) # Combine into a datetime column df['date'] = pd.to_datetime(df[['agno', 'mes', 'dia']]) # Optional: Drop the original year/month/day columns if you don't need them anymore df.drop(['agno', 'mes', 'dia'], axis=1, inplace=True)
Now you have a date column that's a Pandas datetime64 type, which lets you do cool stuff like filter data by month, calculate monthly average temperatures, or plot trends over time.
Manually loading each file is tedious—let's automate it with glob to load and combine all 10 datasets into one big DataFrame:
import glob # Get paths to all your dataset files (adjust the path pattern to match your files) file_paths = glob.glob("path/to/your/datasets/*.csv") # Initialize an empty list to store each processed DataFrame processed_dfs = [] for file in file_paths: # Load the file temp_df = pd.read_csv(file) # Clean column names temp_df.columns = temp_df.columns.str.replace(" ", "_").str.lower() # Merge date columns temp_df['date'] = pd.to_datetime(temp_df[['agno', 'mes', 'dia']]) # Add to our list processed_dfs.append(temp_df) # Combine all datasets into one combined_df = pd.concat(processed_dfs, ignore_index=True) # Check out the combined dataset's info to verify combined_df.info()
This will give you a single DataFrame with all your temperature data, ready for analysis.
df1.info() Output For context, here's how your sample df1.info() might look formatted properly (I filled in the missing bits based on your description):
<class 'pandas.core.frame.DataFrame'> RangeIndex: 17301 entries, 0 to 17300 Data columns (total 4 columns): agno 17301 non-null int64 mes 17301 non-null int64 dia 17301 non-null int64 temp 17301 non-null float64 dtypes: float64(1), int64(3) memory usage: 540.8 KB
If your actual column names have spaces (like "Daily Temp"), the column cleaning step above will handle that seamlessly.
内容的提问来源于stack exchange,提问作者Sergio

