Pandas多级索引选择困惑及新增计算行需求咨询
No worries, MultiIndex can feel tricky at first, but once you get the hang of slicing and index manipulation, it's straightforward. Let's walk through exactly how to create those 'CH' rows you need using your sample data.
First, let's make sure we have your data set up correctly as a MultiIndex DataFrame. Here's the code to replicate it:
import pandas as pd import numpy as np # Build the MultiIndex with the three levels: first, second, third index = pd.MultiIndex.from_tuples( [ ("C", "one", "mean"), ("C", "one", "std"), ("C", "two", "mean"), ("C", "two", "std"), ("C", "three", "mean"), ("C", "three", "std"), ("H", "one", "mean"), ("H", "one", "std"), ("H", "two", "mean"), ("H", "two", "std"), ("H", "three", "mean"), ("H", "three", "std"), ("V", "one", "mean"), ("V", "one", "std"), ("V", "two", "mean"), ("V", "two", "std"), ("V", "three", "mean"), ("V", "three", "std"), ], names=["first", "second", "third"] ) # Populate the data values data = [ [3,4,2,7], [4,1,7,7], [3,1,4,7], [5,6,7,0], [7,0,2,5], [7,3,7,1], [2,4,3,3], [5,5,3,5], [5,7,0,6], [0,1,0,2], [5,2,5,1], [9,0,4,6], [3,7,3,9], [8,7,9,3], [1,9,9,0], [1,1,5,1], [3,1,0,6], [6,2,7,4], ] # Create the DataFrame df = pd.DataFrame(data, index=index, columns=[0,1,2,3])
Step 1: Calculate 'CH' Mean Rows
We need to subtract the 'H' mean values from the 'C' mean values for each 'second' level (one, two, three).
Use slice(None) to select all values in a MultiIndex level. Here's how to slice the mean rows for C and H:
# Slice all mean rows for C and H c_mean = df.loc[("C", slice(None), "mean"), :] h_mean = df.loc[("H", slice(None), "mean"), :] # Calculate the difference: C mean - H mean ch_mean = c_mean.subtract(h_mean) # Update the 'first' level of the index to 'CH' ch_mean.index = ch_mean.index.set_levels(["CH"], level="first")
Step 2: Calculate 'CH' Std Rows
From your partial instruction, I assume you want the standard deviation of the difference between C and H values. For independent variables, the variance of the difference is the sum of variances, so we'll compute the square root of (C_std² + H_std²):
# Slice all std rows for C and H c_std = df.loc[("C", slice(None), "std"), :] h_std = df.loc[("H", slice(None), "std"), :] # Calculate variance sum, then take square root for std ch_var = c_std.pow(2).add(h_std.pow(2)) ch_std = np.sqrt(ch_var) # Update the 'first' level to 'CH' ch_std.index = ch_std.index.set_levels(["CH"], level="first")
Step 3: Combine and Append to Original DataFrame
Now we'll combine the CH mean and std rows, sort them to match the original index structure, and append to the original DataFrame:
# Combine CH mean and std into one DataFrame ch_df = pd.concat([ch_mean, ch_std]).sort_index() # Append to original DataFrame and sort the full index final_df = pd.concat([df, ch_df]).sort_index()
Final Result Preview
If you print final_df.loc["CH"], you'll see the new rows:
0 1 2 3 second one mean 1.000000 0.000000 -1.000000 4.000000 std 6.403124 5.099020 7.615773 8.602325 two mean -2.000000 -6.000000 4.000000 1.000000 std 5.000000 6.082763 7.000000 2.000000 three mean 2.000000 -2.000000 -3.000000 4.000000 std 11.401754 3.000000 8.062258 6.082763
Key takeaways here:
- Use
slice(None)to select all entries in a MultiIndex level - Use
set_levels()to update specific index levels without changing the rest concat()andsort_index()help keep your MultiIndex organized after adding new rows
内容的提问来源于stack exchange,提问作者Guido

