如何以Pythonic无循环方式查询Pandas多层索引Series?
多层索引Series的批量查询方法
给定如下多层索引的Pandas Series:
import pandas as pd s = pd.Series([1, 2, 3, 4, 5, 6], index=pd.MultiIndex.from_product([["A", "B"], ["c", "d", "e"]]))
输出为:
A c 1 d 2 e 3 B c 4 d 5 e 6 dtype: int64
现在需要根据列表l1 = ['A', 'B', 'A']和l2 = ['c', 'd', 'c']的对应元素作为索引对,批量查询Series的对应值,避免使用循环写法,推荐以下几种Pythonic实现方式:
方法1:构造查询索引后用loc直接取值
利用pd.MultiIndex.from_arrays将两个列表组合成匹配的多层索引,再通过loc批量获取对应值:
l1 = ['A', 'B', 'A'] l2 = ['c', 'd', 'c'] query_index = pd.MultiIndex.from_arrays([l1, l2]) result = s.loc[query_index] # 若需要转为列表格式,可使用 result.tolist()
执行后result为:
A c 1 B d 5 A c 1 dtype: int64
方法2:通过get_indexer获取位置索引后用iloc取值
先获取目标索引在原Series索引中的位置,再通过位置批量取值:
pos_indices = s.index.get_indexer(pd.MultiIndex.from_arrays([l1, l2])) result = s.iloc[pos_indices].tolist()
最终result为列表[1, 5, 1]。
方法3:使用reindex方法批量匹配
reindex会根据传入的索引重新排列原Series,自动匹配对应值:
result = s.reindex(pd.MultiIndex.from_arrays([l1, l2])).tolist()
同样得到列表[1, 5, 1]。
以上方法均采用Pandas的矢量化操作,避免了显式循环,代码更简洁且执行效率更高。
内容的提问来源于stack exchange,提问作者user2743931
相关产品推荐
相关产品推荐

