如何将DataFrame指定分组的行添加至其他分组并修改单列值
实现方案说明
完全可以通过merge操作简便实现该需求,无需复杂的循环或分组拼接操作,核心逻辑是通过笛卡尔积合并生成待追加的Circle复制行,再和原数据拼接即可。
核心实现代码
首先构造测试用的输入DataFrame:
import pandas as pd # 构造示例输入 data = { "Type": ["Circle", "Circle", "Circle", "Square", "Square", "Triangle", "Triangle"], "Color": ["Blue", "Green", "Black", "Red", "White", "Red", "White"], "Size": ["Large", "Small", "Large", "Large", "Small", "Large", "Small"], "Dimensions": ["2D", "3D", "2D", "2D", "3D", "2D", "3D"] } df = pd.DataFrame(data)
方案1:cross merge 写法(pandas 1.2.0及以上版本支持,最简洁)
# 提取所有Circle行,删除Type字段后续替换 circle_df = df[df["Type"] == "Circle"].drop(columns="Type") # 提取除Circle外的其他唯一Type值 other_type_df = pd.DataFrame({"Type": df["Type"].loc[df["Type"] != "Circle"].unique()}) # 笛卡尔积合并,生成所有要追加的行 append_df = other_type_df.merge(circle_df, how="cross") # 拼接原始数据和追加行,按Type排序得到最终结果 result = pd.concat([df, append_df], ignore_index=True).sort_values(by="Type", ignore_index=True)
方案2:兼容低版本pandas的写法
circle_df = df[df["Type"] == "Circle"].drop(columns="Type") other_types = df["Type"].loc[df["Type"] != "Circle"].unique() # 用MultiIndex构造笛卡尔积关联关系 append_df = pd.MultiIndex.from_product( [other_types, circle_df.index], names=["Type", "tmp_idx"] ).to_frame(index=False) append_df = append_df.merge(circle_df, left_on="tmp_idx", right_index=True).drop(columns="tmp_idx") result = pd.concat([df, append_df], ignore_index=True).sort_values(by="Type", ignore_index=True)
两种方案输出结果完全匹配预期输出,全程使用pandas向量化操作,数据量越大相比循环写法的性能优势越明显。
内容的提问来源于stack exchange,提问作者Laskii
相关产品推荐
相关产品推荐

