如何利用for循环与range为Pandas DataFrame按年份命名并避免数据覆盖
Hey there! I get that you want to preserve each year's processed DataFrame instead of overwriting them in your loop—this is a super common scenario when working with time-series data like your Nord Pool spot prices. Let's walk through a couple of practical solutions, starting with the most recommended approach.
1. Use a Dictionary to Store Yearly DataFrames (Recommended)
Instead of creating separate named variables (like df_long_2015), using a dictionary is cleaner, easier to manage, and makes merging all your data later a breeze. Here's how to adjust your code:
import pandas as pd # Initialize an empty dictionary to hold each year's DataFrame yearly_dfs = {} for aar in range(2015, 2021 + 1): print(aar) url = f'https://www.nordpoolgroup.com/48c8e5/globalassets/marketdata-excel-files/elspot-prices_{aar}_daily_nok.xls' liste = pd.read_html(url, parse_dates=True, decimal=',', thousands='.', header=2, index_col=0, encoding='UTF-8') df = pd.DataFrame(liste[0]) df.index = pd.to_datetime(df.index, format='%Y-%m-%d') df_long = df.stack().to_frame() df_long.reset_index(inplace=True) df_long.columns = ['Dato','Område','Pris'] filt = df_long['Område'].isin(['Oslo','Bergen','Tr.heim','Tromsø','Kr.sand','Molde']) df_long = df_long.loc[filt, :] # Store the processed DataFrame in the dictionary with the year as key yearly_dfs[aar] = df_long # Or if you prefer the "df_long_2015" style key: # yearly_dfs[f'df_long_{aar}'] = df_long
Accessing Individual Year DataFrames
You can pull out any year's data by referencing the dictionary key:
# Get 2015 data df_2015 = yearly_dfs[2015] # Or if you used the string key: # df_2015 = yearly_dfs['df_long_2015']
Merging All Yearly DataFrames
Once you have all data in the dictionary, merging into a single DataFrame is straightforward:
combined_df = pd.concat(yearly_dfs.values(), ignore_index=True)
2. Create Separate Named Variables (Not Recommended)
If you really need individual variables like df_long_2015, you can use Python's globals() function to dynamically create them. However, this approach is less clean—too many loose variables can make your code harder to debug and maintain. Here's how it would look:
import pandas as pd for aar in range(2015, 2021 + 1): print(aar) url = f'https://www.nordpoolgroup.com/48c8e5/globalassets/marketdata-excel-files/elspot-prices_{aar}_daily_nok.xls' liste = pd.read_html(url, parse_dates=True, decimal=',', thousands='.', header=2, index_col=0, encoding='UTF-8') df = pd.DataFrame(liste[0]) df.index = pd.to_datetime(df.index, format='%Y-%m-%d') df_long = df.stack().to_frame() df_long.reset_index(inplace=True) df_long.columns = ['Dato','Område','Pris'] filt = df_long['Område'].isin(['Oslo','Bergen','Tr.heim','Tromsø','Kr.sand','Molde']) df_long = df_long.loc[filt, :] # Dynamically create a variable named df_long_YYYY globals()[f'df_long_{aar}'] = df_long
Why This Is Less Ideal
- You’ll have to manually track all variable names (e.g.,
df_long_2015,df_long_2016) instead of accessing them from a single dictionary. - It’s easier to introduce bugs if you mistype variable names later on.
Why Your Original Code Had the Overwrite Issue
In your original loop, you’re reassigning the same df_long variable each time the loop runs. Python replaces the old value with the new one in each iteration, which is why you only end up with the last year’s data. Storing each iteration’s result in a dictionary (or as separate variables) preserves each DataFrame independently.
内容的提问来源于stack exchange,提问作者Gisle_B

