如何用Pandas筛选DataFrame列并添加索引与单位参考行?
Great question! Let's break this down into two parts: optimizing your column selection, and adding that unit row (or setting up a proper multi-level column index, which is usually a cleaner approach for this scenario).
Step 1: Optimize Column Selection
Your current column filtering code works, but it's unnecessarily verbose. You can skip the nested loops entirely by leveraging the dictionary's keys directly. To make it safe (in case some keys don't exist in your DataFrame's columns), use a list comprehension to filter only existing columns:
# Rename your dictionary to avoid shadowing Python's built-in `dict` type col_dict = {'time': 'date', 'place': 'London'} # Filter columns that exist in df and are present in the dictionary selected_cols = [col for col in col_dict.keys() if col in df.columns] final_df = df[selected_cols].copy()
If you're certain all keys in the dictionary exist as columns in df, you can simplify even further:
final_df = df[list(col_dict.keys())].copy()
Step 2: Add Unit Metadata
There are two common approaches here, depending on your exact needs:
Option 1: Use a Multi-Level Column Index (Recommended)
This is the most pandas-idiomatic approach because units are metadata about your columns, not actual data rows. You’ll create a two-level column index where the first level is the original column name (dictionary key) and the second level is the corresponding unit (dictionary value):
# Create tuples pairing each selected column with its unit multi_col_tuples = [(col, col_dict[col]) for col in final_df.columns] # Set the multi-index for columns final_df.columns = pd.MultiIndex.from_tuples(multi_col_tuples, names=['Column', 'Unit'])
Your final DataFrame will look like this (example):
time place date London 0 2023-01-01 UK 1 2023-01-02 Paris 2 2023-01-03 London
Option 2: Insert Keys and Units as Actual Data Rows
If you specifically need these to be the first two rows of your DataFrame (instead of column metadata), create temporary Series for the key row and unit row, then concatenate them with your filtered data:
# Create the row with dictionary keys key_row = pd.Series(final_df.columns, index=final_df.columns, name='Index Row') # Create the row with corresponding units unit_row = pd.Series([col_dict[col] for col in final_df.columns], index=final_df.columns, name='Unit Row') # Combine the rows with your original data (ignore_index resets the row index) final_df = pd.concat([key_row.to_frame().T, unit_row.to_frame().T, final_df], ignore_index=True)
This will result in a DataFrame where the first row is the column names, the second is the units, and the rest are your original data entries.
内容的提问来源于stack exchange,提问作者srgam

