You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python Pandas如何避免行遍历高效构建指定结构业务字典

报错原因

  1. 第一个报错是你错误将过滤条件、列名组成的列表直接赋值给了temp_df,列表本身没有pandas的loc方法,调用必然失败。
  2. 第二个报错是字典的key必须是可哈希的不可变对象,你直接传入pandas整列数据(Series类型)作为key,就会触发不可哈希的错误。

优化实现方案

完全避免原始数据层的行遍历,改用pandas原生的分组聚合能力,处理效率比iterrows遍历全量数据高数十到上百倍,完整可运行代码如下:

import pandas as pd
from collections import defaultdict

# 限额数值转换函数
def parse_limit(x):
    if pd.isna(x):
        return pd.NA
    x_str = str(x).strip()
    try:
        # 提取首个数字片段,移除千位分隔符逗号后转整数
        num_part = x_str.split()[0].replace(',', '')
        return int(num_part)
    except (ValueError, IndexError):
        return pd.NA

# 1. 过滤目标数据
ifeu_df = df[df['Venue MIC'] == 'IFEU'].copy()

# 2. 转换限额列数值格式
ifeu_df['spot_limit'] = ifeu_df['Spot month single limit#'].apply(parse_limit)
ifeu_df['other_limit'] = ifeu_df['Other month limit#'].apply(parse_limit)

# 3. 分组聚合数据
# 聚合每个principal对应的所有产品code为集合
code_agg = ifeu_df.groupby('Principal Venue Product Code')['Venue Product Codes'].agg(set).reset_index(name='code_set')
# 聚合每个principal对应的限额数值(取首个非空值即可)
limit_agg = ifeu_df.groupby('Principal Venue Product Code')[['spot_limit', 'other_limit']].first().reset_index()

# 4. 合并聚合结果并生成目标字典
agg_df = pd.merge(code_agg, limit_agg, on='Principal Venue Product Code')
ifeu_dict = {
    row['Principal Venue Product Code']: [row['spot_limit'], row['other_limit'], row['code_set']]
    for _, row in agg_df.iterrows()
}

输出验证

用你提供的样例数据运行后,ifeu_dict['ATW']的输出为:

[5550, 38800, {'ATH', 'ATC', 'RDF', 'ATQ', 'RCA', 'ATD', 'RCO', 'RDG', 'RDE', 'ATW'}]

和你给出的预期结果完全一致。


内容的提问来源于stack exchange,提问作者Bamir

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 03:36:01