多级索引DataFrame处理:删NaN行,留0级索引并单独重置各1级索引
Got it, let's break down how to solve this problem exactly as you described—first cleaning up NaN rows while keeping your Level 0 index intact, then resetting the Level 1 index to start from 0 for each Level 0 group.
Step 1: Remove Rows with NaN Values
First, we'll use dropna() to eliminate any rows containing NaN values. This preserves your original multi-index structure (including both Level 0 and Level 1) while only removing the problematic rows:
import pandas as pd # Assume your DataFrame is named `df` df_clean = df.dropna()
Step 2: Reset Level 1 Index for Each Level 0 Group
The key here is to use groupby() on the Level 0 index, then generate a new sequential index starting at 0 for each group. The most efficient way to do this is with cumcount(), which creates a counter for each group:
# Remove the original Level 1 index temporarily df_temp = df_clean.reset_index(level=1, drop=True) # Generate a new Level 1 index (starting at 0) for each Level 0 group new_level1 = df_temp.groupby(level=0).cumcount() # Attach the new Level 1 index back to the DataFrame df_final = df_temp.set_index(new_level1, append=True)
Let's Test with an Example
Let's create a sample DataFrame to see this in action:
import numpy as np # Create a multi-index DataFrame with NaN values arrays = [ ['A', 'A', 'A', 'B', 'B', 'C'], [0, 1, 2, 0, 1, 0] ] df = pd.DataFrame({'value': [1, np.nan, 3, 4, 5, np.nan]}, index=arrays)
Original df:
value A 0 1.0 1 NaN 2 3.0 B 0 4.0 1 5.0 C 0 NaN
After dropna() (df_clean):
value A 0 1.0 2 3.0 B 0 4.0 1 5.0
Final result (df_final):
value A 0 1.0 1 3.0 B 0 4.0 1 5.0
Perfect—each Level 0 index ('A' and 'B') now has a Level 1 index starting at 0, and we kept the original Level 0 structure intact.
内容的提问来源于stack exchange,提问作者KOB

