You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

重复Custid记录的均值计算与Python代码实现咨询

问题背景

输入数据表(Custid存在重复记录)

Custidtransactionspricerevenue
1245623002148300
12456520029150800
7890130003090000
676512520024124800
676512740018133200

需求说明

  • 去除Custid的重复记录
  • 对拥有多条记录的Custid,计算transactions和price字段的平均值
  • 重新计算revenue,公式为:transactions平均值 × price平均值

预期输出表

Custidtransactionspricerevenue
1245637502593750
7890130003090000
676512630021132300

原代码问题分析

原代码存在多处语法与逻辑错误:

  1. 读取CSV时语法错误:df1 .pd.read_csv多了空格,且变量名、字符串未加引号
  2. 列名引用错误:df1[custid]未给列名加引号,应为df1['Custid']
  3. 逻辑偏差:用value_counts()仅能得到计数,未按Custid分组计算均值,反而误用全局均值
  4. np.where参数不完整,无法实现按分组处理的逻辑

正确解决方案

使用Pandas的groupby按Custid分组,计算指定字段均值后重新生成revenue列:

import pandas as pd

# 读取数据(确保文件路径正确)
df1 = pd.read_csv("file1.csv")

# 按Custid分组,计算transactions和price的平均值
grouped_df = df1.groupby('Custid')[['transactions', 'price']].mean().reset_index()

# 重新计算revenue
grouped_df['revenue'] = grouped_df['transactions'] * grouped_df['price']

# 查看结果
print(grouped_df)

代码说明

  • groupby('Custid'):按客户ID对数据分组
  • [['transactions', 'price']].mean():对分组后的交易数、价格字段分别计算平均值
  • reset_index():将分组后的索引转换为普通列,恢复Custid为数据列
  • 最后通过两列相乘得到新的revenue值,完全匹配需求

内容的提问来源于stack exchange,提问作者user3369545

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 03:45:40