You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何利用for循环与range为Pandas DataFrame按年份命名并避免数据覆盖

Solution for Saving Yearly Pandas DataFrames Without Overwriting

Hey there! I get that you want to preserve each year's processed DataFrame instead of overwriting them in your loop—this is a super common scenario when working with time-series data like your Nord Pool spot prices. Let's walk through a couple of practical solutions, starting with the most recommended approach.

Instead of creating separate named variables (like df_long_2015), using a dictionary is cleaner, easier to manage, and makes merging all your data later a breeze. Here's how to adjust your code:

import pandas as pd

# Initialize an empty dictionary to hold each year's DataFrame
yearly_dfs = {}

for aar in range(2015, 2021 + 1):
    print(aar)
    url = f'https://www.nordpoolgroup.com/48c8e5/globalassets/marketdata-excel-files/elspot-prices_{aar}_daily_nok.xls'
    liste = pd.read_html(url, parse_dates=True, decimal=',', thousands='.', header=2, index_col=0, encoding='UTF-8')
    df = pd.DataFrame(liste[0])
    df.index = pd.to_datetime(df.index, format='%Y-%m-%d')
    df_long = df.stack().to_frame()
    df_long.reset_index(inplace=True)
    df_long.columns = ['Dato','Område','Pris']
    filt = df_long['Område'].isin(['Oslo','Bergen','Tr.heim','Tromsø','Kr.sand','Molde'])
    df_long = df_long.loc[filt, :]
    
    # Store the processed DataFrame in the dictionary with the year as key
    yearly_dfs[aar] = df_long
    # Or if you prefer the "df_long_2015" style key:
    # yearly_dfs[f'df_long_{aar}'] = df_long

Accessing Individual Year DataFrames

You can pull out any year's data by referencing the dictionary key:

# Get 2015 data
df_2015 = yearly_dfs[2015]
# Or if you used the string key:
# df_2015 = yearly_dfs['df_long_2015']

Merging All Yearly DataFrames

Once you have all data in the dictionary, merging into a single DataFrame is straightforward:

combined_df = pd.concat(yearly_dfs.values(), ignore_index=True)

If you really need individual variables like df_long_2015, you can use Python's globals() function to dynamically create them. However, this approach is less clean—too many loose variables can make your code harder to debug and maintain. Here's how it would look:

import pandas as pd

for aar in range(2015, 2021 + 1):
    print(aar)
    url = f'https://www.nordpoolgroup.com/48c8e5/globalassets/marketdata-excel-files/elspot-prices_{aar}_daily_nok.xls'
    liste = pd.read_html(url, parse_dates=True, decimal=',', thousands='.', header=2, index_col=0, encoding='UTF-8')
    df = pd.DataFrame(liste[0])
    df.index = pd.to_datetime(df.index, format='%Y-%m-%d')
    df_long = df.stack().to_frame()
    df_long.reset_index(inplace=True)
    df_long.columns = ['Dato','Område','Pris']
    filt = df_long['Område'].isin(['Oslo','Bergen','Tr.heim','Tromsø','Kr.sand','Molde'])
    df_long = df_long.loc[filt, :]
    
    # Dynamically create a variable named df_long_YYYY
    globals()[f'df_long_{aar}'] = df_long

Why This Is Less Ideal

  • You’ll have to manually track all variable names (e.g., df_long_2015, df_long_2016) instead of accessing them from a single dictionary.
  • It’s easier to introduce bugs if you mistype variable names later on.

Why Your Original Code Had the Overwrite Issue

In your original loop, you’re reassigning the same df_long variable each time the loop runs. Python replaces the old value with the new one in each iteration, which is why you only end up with the last year’s data. Storing each iteration’s result in a dictionary (or as separate variables) preserves each DataFrame independently.

内容的提问来源于stack exchange,提问作者Gisle_B

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 09:32:40