Pandas多索引Series添加行并归入对应外层索引组(禁用sort_index)
向多索引Series中添加行并归入对应外层索引组(禁用
sort_index()) 需求:给多索引Series新增一行,需将该行归入指定的外层索引组,且必须保留原有索引的字母顺序,因此不能使用df.sort_index()。
原代码
import pandas as pd import numpy as np categories = {"A":["c", "b", "a"] , "B": ["a", "b", "c"], "C": ["a", "b", "d"] } array = [] expected_fields = [] for key, value in categories.items(): array.extend([key]* len(value)) expected_fields.extend(value) arrays = [array ,expected_fields] tuples = list(zip(*arrays)) index = pd.MultiIndex.from_tuples(tuples) df = pd.Series(np.random.randn(9), index=index) df["A", "d"] = 2 print(df)
当前输出
A c 0.887137 b -0.105262 a -0.180093 B a -0.687134 b -1.120895 c 2.398962 C a -2.226126 b -0.203238 d 0.036068 A d 2.000000 <------------ dtype: float64
期望输出
A c 0.887137 b -0.105262 a -0.180093 d 2.000000 <-------------- B a -0.687134 b -1.120895 c 2.398962 C a -2.226126 b -0.203238 d 0.036068 dtype: float64
解决方案
直接通过索引赋值会把新行追加到Series末尾,要让新行归入对应外层索引组,可通过以下两种方式实现:
方法1:拆分原Series后合并新行
import pandas as pd import numpy as np categories = {"A":["c", "b", "a"] , "B": ["a", "b", "c"], "C": ["a", "b", "d"] } array = [] expected_fields = [] for key, value in categories.items(): array.extend([key]* len(value)) expected_fields.extend(value) arrays = [array ,expected_fields] tuples = list(zip(*arrays)) index = pd.MultiIndex.from_tuples(tuples) df = pd.Series(np.random.randn(9), index=index) # 创建要添加的新行 new_row = pd.Series([2], index=pd.MultiIndex.from_tuples([("A", "d")])) # 拆分原Series为A组和其他组 a_group = df.loc["A"] other_groups = df.drop("A") # 合并A组与新行,再合并其他组 updated_a = pd.concat([a_group, new_row]) result = pd.concat([updated_a, other_groups]) print(result)
方法2:通过重新索引插入指定位置
如果需要精确控制新行在组内的位置(比如不是组末),可以先调整索引列表再重新索引:
import pandas as pd import numpy as np categories = {"A":["c", "b", "a"] , "B": ["a", "b", "c"], "C": ["a", "b", "d"] } array = [] expected_fields = [] for key, value in categories.items(): array.extend([key]* len(value)) expected_fields.extend(value) arrays = [array ,expected_fields] tuples = list(zip(*arrays)) index = pd.MultiIndex.from_tuples(tuples) df = pd.Series(np.random.randn(9), index=index) # 获取原索引的列表形式 index_list = list(df.index) # 找到A组最后一个元素的位置 last_a_pos = [i for i, idx in enumerate(index_list) if idx[0] == "A"][-1] # 在A组末尾插入新索引 index_list.insert(last_a_pos + 1, ("A", "d")) # 重新索引并赋值新行 df = df.reindex(index_list) df.loc[("A", "d")] = 2 print(df)
以上两种方法都能保留原有索引的顺序,同时将新行归入指定的外层索引组。
内容的提问来源于stack exchange,提问作者saul santos
相关产品推荐
相关产品推荐

