如何按条件删除Pandas值并相应左移行内排名值
Solution for Shifting Rows When Status is "out"
Hey there! Let's work through this problem together. First, I'll start with a sample DataFrame that aligns with your description—this will make it easier to see the solution in action.
Sample Input DataFrame
Let's create a representative dataset:
import pandas as pd data = { 'status': ['in', 'out', 'in', 'out'], 'A': [10, 20, 30, 40], 'rank1': [1, 2, 3, 4], 'rank2': [5, 6, 7, 8], 'rank3': [9, 10, 11, 12] } df = pd.DataFrame(data)
This gives us:
| status | A | rank1 | rank2 | rank3 |
|---|---|---|---|---|
| in | 10 | 1 | 5 | 9 |
| out | 20 | 2 | 6 | 10 |
| in | 30 | 3 | 7 | 11 |
| out | 40 | 4 | 8 | 12 |
Method 1: Efficient Vectorized Operation (Best for Large Data)
If you're working with a large DataFrame, vectorized operations are way faster than row-wise loops. Here's how to directly update the relevant rows:
# Identify rows where status is 'out' out_mask = df['status'] == 'out' # Shift values left: A gets rank1, rank1 gets rank2, rank2 gets rank3, rank3 becomes NA df.loc[out_mask, 'A'] = df.loc[out_mask, 'rank1'] df.loc[out_mask, 'rank1'] = df.loc[out_mask, 'rank2'] df.loc[out_mask, 'rank2'] = df.loc[out_mask, 'rank3'] df.loc[out_mask, 'rank3'] = pd.NA
Method 2: Flexible Row-wise Apply (Good for Dynamic Column Counts)
If you have a variable number of rank columns and don't want to hardcode each shift, use apply to handle rows dynamically:
def process_row(row): if row['status'] == 'out': # Extract non-status columns, drop 'A', then shift left and pad with NA non_status_cols = row.drop('status') shifted_vals = non_status_cols.drop('A').tolist() + [pd.NA] # Reconstruct the row with status first, then shifted values return pd.Series([row['status']] + shifted_vals, index=row.index) return row result_df = df.apply(process_row, axis=1)
Expected Output
Either method will give you this result:
| status | A | rank1 | rank2 | rank3 |
|---|---|---|---|---|
| in | 10 | 1 | 5 | 9 |
| out | 2 | 6 | 10 | |
| in | 30 | 3 | 7 | 11 |
| out | 4 | 8 | 12 |
Notes
- If your DataFrame has more
rankcolumns, Method 2 will automatically handle the shift without needing to update code. - Use Method 1 for performance if your dataset is large, as vectorized operations are optimized in pandas.
内容的提问来源于stack exchange,提问作者wflwo
相关产品推荐
相关产品推荐

