为时间膨胀多变量Dataframe创建保留各项目初始比率的新列
为多变量Dataframe添加初始比率列
初始Dataframe数据
产品名称表
| product1 | product2 | product3 | product4 | product5 |
|---|---|---|---|---|
| straws | orange | melon | chair | bread |
| melon | milk | book | coffee | cake |
| bread | melon | coffe | chair | book |
计数值表
| CountProduct1 | CountProduct2 | CountProduct3 | Countproduct4 | Countproduct5 |
|---|---|---|---|---|
| 1 | 1 | 1 | 1 | 1 |
| 2 | 1 | 1 | 1 | 1 |
| 2 | 3 | 2 | 2 | 2 |
比率值表
| RatioProduct1 | RatioProduct2 | RatioProduct3 | Ratioproduct4 | Ratioproduct5 |
|---|---|---|---|---|
| 0.28 | 0.54 | 0.33 | 0.35 | 0.11 |
| 0.67 | 0.25 | 0.13 | 0.11 | 0.59 |
| 2.5 | 1.69 | 1.9 | 2.5 | 1.52 |
需求说明
需要新增五列InitialRatio1至InitialRatio5,每列对应原表中同位置产品在整个Dataframe里的初始比率(即该产品第一次出现时对应的比率值)。预期输出如下:
预期输出(仅新增列)
| InitialRatio1 | InitialRatio2 | InitialRatio3 | InitialRatio4 | InitialRatio5 |
|---|---|---|---|---|
| 0.28 | 0.54 | 0.33 | 0.35 | 0.11 |
| 0.33 | 0.25 | 0.13 | 0.31 | 0.59 |
| 0.11 | 0.33 | 0.31 | 0.35 | 0.13 |
解决步骤(基于Pandas)
1. 合并初始数据为单个Dataframe
先把三个表格合并,方便后续统一处理:
import pandas as pd # 构造产品名称Dataframe products_df = pd.DataFrame( [["straws", "orange", "melon", "chair", "bread"], ["melon", "milk", "book", "coffee", "cake"], ["bread", "melon", "coffe", "chair", "book"]], columns=["product1", "product2", "product3", "product4", "product5"] ) # 构造计数值Dataframe counts_df = pd.DataFrame( [[1,1,1,1,1], [2,1,1,1,1], [2,3,2,2,2]], columns=["CountProduct1", "CountProduct2", "CountProduct3", "Countproduct4", "Countproduct5"] ) # 构造比率值Dataframe ratios_df = pd.DataFrame( [[0.28,0.54,0.33,0.35,0.11], [0.67,0.25,0.13,0.11,0.59], [2.5,1.69,1.9,2.5,1.52]], columns=["RatioProduct1", "RatioProduct2", "RatioProduct3", "Ratioproduct4", "Ratioproduct5"] ) # 合并三个Dataframe df = pd.concat([products_df, counts_df, ratios_df], axis=1)
2. 构建产品-初始比率映射字典
遍历所有产品列,记录每个产品第一次出现时对应的比率值:
product_initial_ratio = {} # 遍历每一组产品列和对应比率列 for i in range(1, 6): product_col = f"product{i}" # 注意原比率列名的大小写不一致,单独处理第4列 ratio_col = f"RatioProduct{i}" if i !=4 else "Ratioproduct4" for product, ratio in zip(df[product_col], df[ratio_col]): if product not in product_initial_ratio: product_initial_ratio[product] = ratio
3. 生成目标初始比率列
根据映射字典,为每个产品位置匹配对应的初始比率:
# 生成五列InitialRatio for i in range(1, 6): product_col = f"product{i}" new_col = f"InitialRatio{i}" df[new_col] = df[product_col].map(product_initial_ratio) # 查看新增列结果 print(df[["InitialRatio1", "InitialRatio2", "InitialRatio3", "InitialRatio4", "InitialRatio5"]])
运行代码后,新增列的结果将与预期输出完全一致。
内容的提问来源于stack exchange,提问作者Bry Sab
相关产品推荐
相关产品推荐

