解决循环计算fPAR分位数合并DataFrame时的索引报错问题
问题
我尝试按年份遍历各个地区,计算fPAR指标从第1分位数到第99分位数(以5个百分点为步长)的数值,但在循环完成后合并结果到pandas DataFrame时,触发报错:Reindexing only valid with uniquely valued Index objects。
我的代码如下:
import pandas as pd import numpy as np x = pd.read_csv("/DATA.csv") quantiles = np.array([0.01, 0.05, 0.1 , 0.15, 0.2 , 0.25, 0.3 , 0.35, 0.4 , 0.45, 0.5 , 0.55, 0.6 , 0.65, 0.7 , 0.75, 0.8 , 0.85, 0.9 , 0.95, 0.99]) # Loop through every unique state and year results = pd.DataFrame(columns=['NAME', 'year'] + list(quantiles)) # Initialize empty results DataFrame for state in x['NAME'].unique(): for year in x['year'].unique(): state_year_df = x[(x['NAME'] == state) & (x['year'] == year)] # Subset the data for the current state and year if not state_year_df.empty: # Skip empty subsets state_year_quantiles = state_year_df['fPAR'].quantile(q=quantiles) # Calculate the quantiles state_year_results = pd.concat([pd.Series([state, year]), state_year_quantiles.reset_index(drop=True)], axis=0) # Combine state, year, and quantiles into a single Series state_year_results = state_year_results.fillna(0) results = results.append(state_year_results, ignore_index=True) # Append the results for the current state and year to the overall results DataFrame print('Quantiles:', state_year_results)
解决方案
报错核心原因是手动拼接的state_year_results没有与results的列名对应,且已弃用的append方法易引发索引不匹配问题。推荐两种修复方式:
方式一:优化循环逻辑,用列表存储结果后一次性合并
通过构造字典匹配列名,避免索引混乱,同时提升运行效率:
import pandas as pd import numpy as np x = pd.read_csv("/DATA.csv") quantiles = np.array([0.01, 0.05, 0.1, 0.15, 0.2, 0.25, 0.3, 0.35, 0.4, 0.45, 0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, 0.9, 0.95, 0.99]) results_list = [] for state in x['NAME'].unique(): for year in x['year'].unique(): state_year_df = x[(x['NAME'] == state) & (x['year'] == year)] if not state_year_df.empty: # 计算分位数并转为字典,方便匹配列名 quantile_dict = state_year_df['fPAR'].quantile(q=quantiles).to_dict() # 构造当前行的完整数据 row_data = {'NAME': state, 'year': year} row_data.update(quantile_dict) # 存入结果列表 results_list.append(row_data) # 一次性生成最终DataFrame results = pd.DataFrame(results_list).fillna(0) print(results)
方式二:用groupby替代双重循环(更简洁高效)
直接通过分组计算分位数,完全规避索引问题:
import pandas as pd import numpy as np x = pd.read_csv("/DATA.csv") quantiles = np.array([0.01, 0.05, 0.1, 0.15, 0.2, 0.25, 0.3, 0.35, 0.4, 0.45, 0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, 0.9, 0.95, 0.99]) # 按地区和年份分组,直接计算分位数并展开为列 results = x.groupby(['NAME', 'year'])['fPAR'].quantile(q=quantiles).unstack().reset_index() results = results.fillna(0) print(results)
内容的提问来源于stack exchange,提问作者Thomas
相关产品推荐
相关产品推荐

