You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

导入不同商品CSV后,Pandas分组求和及占比结果未更新问题

问题:导入不同CSV后分组计算结果始终一致的排查与解决

我使用私有农场数据库,对应的DataFrame命名为Pome,希望为每种新商品导入对应的CSV文件(原始数据下载前已完成过滤),按'B_Desc'(农场边界描述)列分组,计算各边界的电力使用占总电力的百分比。但导入不同商品的CSV文件后,代码返回的分组求和及总占比结果完全一致。

我的代码如下:

# Split the column into separate columns & new column names
Pome = Pome['Total Input Description;Primary Value;B_Desc;Boundary Value;UOM'].str.split(';', expand=True)
Pome.columns = ['Total Input Description', 'Primary Value', 'B_Desc', 'Boundary Value', 'UOM']


# Replace non-numeric values with NaN
Pome['Boundary Value'] = pd.to_numeric(Pome['Boundary Value'], errors='coerce')
# Replace NaN with a default value (e.g., 0)
Pome['Boundary Value'] = Pome['Boundary Value'].fillna(0)
# Convert to integers
Pome['Boundary Value'] = Pome['Boundary Value'].astype(int)

# Filter the DataFrame to include only rows with 'Total GRID Electricity' in 'Total Input Description'
pome_grid_electricity = Pome[Pome['Total Input Description'].str.contains('Total GRID Electricity', case=False)]

# Group the filtered DataFrame by 'B_Desc' and calculate the sum of 'Boundary Value'
sum_boundary_value = pome_grid_electricity.groupby('B_Desc')['Boundary Value'].sum().reset_index()

# Calculate the total of all groups
total_boundary_value = sum_boundary_value['Boundary Value'].sum()

# Calculate the percentage of each group relative to the total
sum_boundary_value['Percentage of Total'] = (sum_boundary_value['Boundary Value'] / total_boundary_value) * 100

# Rename the 'Boundary Value' column in the resulting DataFrame
sum_boundary_value = sum_boundary_value.rename(columns={'Boundary Value': 'Sum Boundary Value'})

# Print 
print(sum_boundary_value)

解决建议

  • 补全CSV导入逻辑:当前代码缺失读取新CSV的核心步骤,若每次处理新商品时未重新加载文件并覆盖Pome变量,会一直复用初始数据。需在代码最开头添加Pome = pd.read_csv("对应商品的CSV文件路径"),确保每次使用的是新文件的数据。
  • 验证过滤数据的有效性:在过滤后添加print(pome_grid_electricity.shape),查看不同CSV处理后的数据行数是否有差异。若行数完全一致,说明'Total GRID Electricity'的匹配可能存在问题(比如文本拼写、格式不一致),需检查CSV中的对应字段内容,调整过滤条件。
  • 确认列拆分的正确性:拆分后打印Pome.head(),对比不同CSV的列内容是否符合预期。若新CSV的原始列分隔符、列顺序与初始数据不同,会导致拆分后Boundary Value列数据错误,需调整str.split的参数或列名映射逻辑。
  • 清理环境缓存:若在Jupyter Notebook等交互式环境中运行,可能存在旧变量残留问题。可在代码开头执行%reset -f清空变量,或重启内核后重新运行,避免旧数据干扰。

内容的提问来源于stack exchange,提问作者tanacubana

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 01:17:33