Pandas:如何在不排序数据集的情况下选取特定值行之间的所有行
Got it, let's tackle this problem—you need to pull out all rows that lie between specific marker rows in your dataset without sorting (so you keep the original sequence and all critical context). Below is a practical, step-by-step solution using Pandas, which is ideal for handling tabular data like your examples.
Example 1: Dataset with event_type Markers
First, let's work with your first dataset, where markers are rows where event_type is 1 or 2.
Step 1: Load the Data
import pandas as pd # Your first dataset data1 = pd.DataFrame({ 'game_clock': [711, 710, 709, 708, 707, 706], 'quarter': [1, 1, 2, 3, 4, 4], 'event_type': [1, 3, 4, 2, 4, 1] })
Step 2: Identify Marker Rows
First, find the indices of all rows that match your marker condition (event_type is 1 or 2):
marker_indices = data1[data1['event_type'].isin([1, 2])].index.tolist() # Output: [0, 3, 5]
Step 3: Generate Extraction Ranges
We need to create ranges that cover:
- Rows between each pair of consecutive markers (including the markers themselves)
- Any rows before the first marker or after the last marker (if they exist)
ranges = [] # Handle rows before the first marker (if any) if marker_indices[0] > 0: ranges.append((0, marker_indices[0])) # Handle ranges between consecutive markers for i in range(len(marker_indices) - 1): ranges.append((marker_indices[i], marker_indices[i+1])) # Handle rows after the last marker (if any) if marker_indices[-1] < len(data1) - 1: ranges.append((marker_indices[-1], len(data1) - 1))
Step 4: Extract and Combine Rows
# Extract each range and concatenate the results result1 = pd.concat([data1.loc[start:end] for start, end in ranges]) # Optional: Reset index if you don't need to keep the original indices result1 = result1.reset_index(drop=True)
The final result will include:
- Rows 0 to 3 (from the first marker to the second marker)
- Rows 3 to 5 (from the second marker to the third marker)
Example 2: Dataset with A Column Markers
Your second dataset has markers where A is 1 or 2, and you want to extract specific ranges including individual marker rows and segments between them. Here's how to adapt the solution:
Step 1: Load the Data
data2 = pd.DataFrame({ 'A': [1,2,2,4,3,3,3,1,4,3,3,1,3,4,1,4,1,3,4,1,1,1,4,1,2,1,1,1,4,2], 'B': [0.278179,0.069914,0.633110,0.584766,0.581232,0.677205,0.687155,0.438927,0.320927,0.570552,0.479849,0.861074,0.834805,0.105766,0.060408,0.596882,0.792395,0.226356,0.535201,0.136066,0.372244,0.151977,0.429822,0.792706,0.406957,0.177850,0.909252,0.545331,0.100497,0.718721] })
Step 2: Identify Marker Rows
marker_indices = data2[data2['A'].isin([1, 2])].index.tolist() # Output: [0,1,2,7,11,14,19,20,21,24,25,26,27,29]
Step 3: Generate Custom Extraction Ranges
For this dataset, you want to extract individual marker rows and the segments between them. We'll adjust the range logic to handle this:
ranges = [] prev_marker = -1 for idx in marker_indices: # Extract the segment between the last marker and current marker (if there's a gap) if idx > prev_marker + 1: ranges.append((prev_marker + 1, idx)) # Extract the current marker row as a single-row range ranges.append((idx, idx)) prev_marker = idx # Extract any rows after the last marker if prev_marker < len(data2) - 1: ranges.append((prev_marker + 1, len(data2) - 1))
Step 4: Extract and Combine Rows
result2 = pd.concat([data2.loc[start:end] for start, end in ranges]) result2 = result2.reset_index(drop=True)
This result will exactly match the subset you specified:
- Individual marker rows like
[0],[1],[2] - Segments like
[3:8](rows 3-7),[8:12](rows 8-11), and[28:](rows 28-29)
Customization Tips
- Skip marker rows: If you only want the rows between markers (not the markers themselves), adjust the range start/end to
(prev_marker + 1, idx - 1)for segments, and skip adding the single marker rows. - Merge consecutive markers: If you don't need to extract consecutive marker rows individually, group them into a single range by checking if
idx == prev_marker + 1and merging the ranges. - Large datasets: This method is efficient even for big datasets because it works directly with original indices—no sorting or shuffling required.
内容的提问来源于stack exchange,提问作者ThP

