基于列名的Python列计算:DataFrame新增差值列需求
实现DataFrame列间差值计算并插入对应位置
以下是基于列名匹配的解决方案,用Pandas完成需求:
步骤说明
- 分类并排序列:先提取所有
LHA_和JH_开头的列,再按下划线后的数字编号排序,确保配对逻辑的正确性。 - 计算差值并构建新列顺序:从第二个
LHA列开始,对应前一个编号的JH列计算差值,同时按要求的顺序将原列和新差值列加入列顺序列表。 - 重构DataFrame:用新的列顺序重新排列原DataFrame,得到最终结构。
代码示例
import pandas as pd # 示例DataFrame(可替换为你的实际数据) data = { 'LHA_1': [10, 20, 30], 'JH_1': [5, 6, 7], 'LHA_2': [15, 25, 35], 'JH_2': [8, 9, 10], 'LHA_3': [20, 30, 40], 'JH_3': [11, 12, 13], 'LHA_4': [25, 35, 45], 'JH_5': [14, 15, 16] } df = pd.DataFrame(data) # 提取并按编号排序LHA、JH列 lha_cols = sorted([col for col in df.columns if col.startswith('LHA_')], key=lambda x: int(x.split('_')[1])) jh_cols = sorted([col for col in df.columns if col.startswith('JH_')], key=lambda x: int(x.split('_')[1])) # 初始化新列顺序列表 new_col_order = ['LHA_1', 'JH_1'] # 遍历配对计算差值,填充列顺序 for idx in range(1, len(lha_cols)): current_lha = lha_cols[idx] target_jh_idx = idx - 1 # LHA_n对应JH_{n-1} # 若没有对应JH列,直接添加当前LHA列 if target_jh_idx >= len(jh_cols): new_col_order.append(current_lha) continue target_jh = jh_cols[target_jh_idx] # 添加当前LHA列 new_col_order.append(current_lha) # 计算并添加差值列 diff_col_name = f"{current_lha} - {target_jh}" df[diff_col_name] = df[current_lha] - df[target_jh] new_col_order.append(diff_col_name) # 添加对应编号的JH列(如果存在) if idx < len(jh_cols): new_col_order.append(jh_cols[idx]) # 补充未加入的剩余JH列 for jh_col in jh_cols: if jh_col not in new_col_order: new_col_order.append(jh_col) # 按新顺序重构DataFrame df = df[new_col_order] print(df)
运行后输出的DataFrame会完全符合你要求的格式:LHA_1 → JH_1 → LHA_2 → LHA_2-JH_1 → JH_2 → LHA_3 → LHA_3-JH_2 → JH_3 → LHA_4 → JH_5...
内容的提问来源于stack exchange,提问作者Connie Xu
相关产品推荐
相关产品推荐

