Pandas多索引切片问题:移除冗余行列与特征切片报错解决
Answer to Your Pandas CSV Questions
Hey there! As a fellow pandas beginner, I totally get these little hurdles when working with multi-indexes and column slicing. Let's break down your two problems step by step:
Problem 1: Removing id and header after using df.xs('Run')
When you run df.xs('Run'), the id is part of the resulting DataFrame's index (since you set ['Activity', 'id'] as the multi-index for your original df). To remove it and exclude the header row when saving to CSV, here's what you can do:
# Get all Run data run_data = df.xs('Run') # Drop the id index and save without header run_data.reset_index(drop=True).to_csv('run_results.csv', header=False)
reset_index(drop=True): This removes theidindex entirely, leaving only your value columns.header=False: Tells pandas not to write the column names (likevalue,value.1) to the CSV file.
If you just want to view the data without id and header (not saving yet), you can print it like this:
print(run_data.reset_index(drop=True).to_string(header=False))
Problem 2: Correctly slicing columns with loc
The error you're seeing happens because df.loc uses label-based indexing—meaning it looks for column names that match the values you pass. Your columns are named value, value.1, value.2, etc., not integers like 2 or 11.
To fix this, you have two straightforward options:
Option 1: Use iloc for positional slicing
If you know the integer positions of the columns you want, use iloc (which works with positions instead of labels):
idx = pd.IndexSlice # Get the Run, 1 rows, then slice columns from position 2 onwards (adjust as needed) subset = df.loc[idx['Run', 1], :].iloc[:, 2:]
Option 2: Slice using column labels
Since your columns follow a pattern (value, value.1, ...), you can filter or slice the column names directly:
idx = pd.IndexSlice # Get all columns except the first 'value' column cols_to_use = df.columns[1:] # Or get columns from 'value.1' to 'value.8' explicitly cols_to_use = df.columns[1:9] subset = df.loc[idx['Run', 1], cols_to_use]
A quick tip: If you find multi-indexes confusing while starting out, you can keep id as a regular column instead of part of the index by skipping set_index(['Activity', 'id']) or using reset_index() later. This might make column slicing feel more intuitive at first.
内容的提问来源于stack exchange,提问作者Eypros

