如何用Pandas高效将DataFrame子集转换为独立DataFrame
Pandas 多级索引DataFrame结构转换优化问题
我有如下结构的Pandas DataFrame:
Index Key 2010-01 2010-02 2010-03 ... 2020-12 A/B/C foo 0.23 0.44 0 2.1 A/B/C bar 0.43 0.12 0.23 1.2 A/B/C baz 0.25 0.23 0.2 2.5 P/Q/R foo 0.31 0.41 0 2.4 P/Q/R foo 0.33 0.54 0.5 4.2 P/Q/R foo 0.93 0.64 0.99 6.5
其中index是多级索引(multi-column index),每个索引下都存在"foo"、"bar"、"baz"。
我希望将这类数据转换为如下结构的独立DataFrame:
# dataframe for A/B/C index foo bar baz 2010-01 0.23 0.43 0.25 2010-02 0.44 0.12 0.23 ... 2020-12 2.1 1.2 2.5
我刚接触Pandas,尝试过将数据转换为字典处理,已有一个分两步的解决方案,伪代码如下:
# loop over the converted dictionary (as per keys) For each key, create 'foo', 'bar', 'baz' with empty dicts; when encountering a row for 'foo', collect all values from col 2010-01 to 2020-12 as a list do the same for 'bar' and 'baz'. Add to the nested dict that is held by the given key For the second pass, loop through each key take the nested dict and create a dataframe using the entire dict and the dates 2010-01 to 2020-12 as the index.
问题:
有没有更符合Pandas风格的实现方法?能否在转换A/B/C这类分组数据时进行转置,同时避免转置带来的性能损耗?
实际数据包含10000+个此类索引(超过3万行)。
内容的提问来源于stack exchange,提问作者Serendipity
相关产品推荐
相关产品推荐

