如何在R中复制数据框行并修改Agent_Count生成假设数据?
解决方案:扩展预测分析用的假设DataFrame
完整实现代码
import pandas as pd import numpy as np # 构建原始DataFrame data = { 'Group': ['Clinical Support', 'PW Reset', 'Technical Support'], 'Agent_Count': [11.75, 12.06, 21.15], 'HOUR': [9, 9, 9], 'Answered_Calls': [52.69, 53.79, 81.02], 'Aban': [2.77, 22.27, 2.22], 'Total_Calls': [56.65, 81.98, 84.20] } df = pd.DataFrame(data) # 设定每个分组要生成的假设行数(20-30之间调整) num_hypotheses = 25 # 初始化扩展后的空DataFrame expanded_df = pd.DataFrame() # 遍历每个原始分组行,生成假设数据 for _, row in df.iterrows(): # 生成指定数量的Agent_Count假设值(示例:原数值±20%范围内的随机数) agent_variants = np.random.uniform( low=row['Agent_Count'] * 0.8, high=row['Agent_Count'] * 1.2, size=num_hypotheses ) # 保留两位小数,和原始数据格式统一 agent_variants = np.round(agent_variants, 2) # 复制当前行的其他列,生成临时数据片段 temp_segment = pd.DataFrame({ 'Group': [row['Group']] * num_hypotheses, 'Agent_Count': agent_variants, 'HOUR': [row['HOUR']] * num_hypotheses, 'Answered_Calls': [row['Answered_Calls']] * num_hypotheses, 'Aban': [row['Aban']] * num_hypotheses, 'Total_Calls': [row['Total_Calls']] * num_hypotheses }) # 合并到最终结果 expanded_df = pd.concat([expanded_df, temp_segment], ignore_index=True) # 查看结果示例 print(expanded_df.head())
关键细节说明
- 假设值生成方式:示例中用
np.random.uniform在原Agent_Count的80%-120%范围内生成随机值,你可以根据业务需求替换为:- 线性序列:用
np.linspace(start, stop, num=num_hypotheses)生成均匀分布的不重复值 - 固定步长序列:比如
np.arange(row['Agent_Count']-3, row['Agent_Count']+3, 0.2)
- 线性序列:用
- 行数控制:修改
num_hypotheses的值即可调整每个分组生成的假设行数(20-30之间任选) - 数据一致性:所有非Agent_Count的列完全复制原始行的平均值,保证假设数据的其他变量不变,仅调整Agent_Count用于预测分析
内容的提问来源于stack exchange,提问作者PeatyBoWeaty
相关产品推荐
相关产品推荐

