如何在Pandas中为DataFrame新增列以选中指定行区间的行?
How to Add a Column Marking Rows Between Two Specific Rows in Pandas
First, let's start with your sample DataFrame for reference:
import pandas as pd data = {'currency': ['Euro', 'Euro', 'Euro', 'Dollar', 'Dollar', 'Yen', 'Yen', 'Yen', 'Pound', 'Pound', 'Pound', 'Pesos', 'Pesos'], 'cost': [34, 67, 32, 29, 48, 123, 23, 45, 78, 86, 23, 45, 67]} df = pd.DataFrame(data, columns=['currency', 'cost'])
Option 1: Specify Rows by Index
If you know the integer indices of the start and end rows, you can directly slice the DataFrame to mark the range.
Example:
Suppose you want to mark rows between index 3 (Dollar, 29) and index 7 (Yen, 45) inclusive:
# Define your input parameters (start and end indices) start_idx = 3 end_idx = 7 # Initialize the new column with False df['in_range'] = False # Set rows between start and end to True # Use min and max to handle cases where start_idx > end_idx df.loc[min(start_idx, end_idx):max(start_idx, end_idx), 'in_range'] = True
This will set in_range to True for rows 3 through 7.
Option 2: Specify Rows by Column Values
If you don't know the indices, but can identify the start and end rows using column values (e.g., specific currency and cost combinations), first find their indices, then apply the same logic as above.
Example:
Suppose you want to start at the row where currency == 'Dollar' and cost == 29, and end at currency == 'Yen' and cost == 45:
# Find start and end indices based on conditions start_idx = df[(df['currency'] == 'Dollar') & (df['cost'] == 29)].index[0] end_idx = df[(df['currency'] == 'Yen') & (df['cost'] == 45)].index[0] # Create the in_range column df['in_range'] = False df.loc[min(start_idx, end_idx):max(start_idx, end_idx), 'in_range'] = True
Useful Notes:
- The
minandmaxensure the code works even if you accidentally pass the start index as larger than the end index. - If there are multiple rows matching your condition (e.g., multiple 'Dollar' rows with cost 29),
index[0]picks the first match. Adjust this toindex[-1]if you need the last matching row instead. - To make this reusable, wrap it in a function:
def mark_row_range(df, start_condition, end_condition): start_idx = df[start_condition].index[0] end_idx = df[end_condition].index[0] df['in_range'] = False df.loc[min(start_idx, end_idx):max(start_idx, end_idx), 'in_range'] = True return df # Usage example: df = mark_row_range(df, start_condition=(df['currency'] == 'Dollar') & (df['cost'] == 29), end_condition=(df['currency'] == 'Yen') & (df['cost'] == 45))
内容的提问来源于stack exchange,提问作者Dsto4
相关产品推荐
相关产品推荐

