如何筛选Pandas DataFrame中数值区间首次出现的行?
Solution to Filter First Continuous Block of Rows in Pandas
Got it, let's fix this up. Your current code grabs all rows where x falls in (2,6], but you only want the first continuous block of such rows (ignoring any later entries that re-enter the interval). Here's how to adjust your code to make that happen:
Full Working Code
import pandas as pd # Your original data data = {'x':[1,2,3,4,5,6,7,6.5,5.5,4.5,3.5,2.5,1], 'y':[1,4,3,3,52,3,74,64,15,41,31,12,11]} df = pd.DataFrame(data) xlim = [2,6] # Original mask to identify rows in the target range mask = (df['x'] > xlim[0]) & (df['x'] <= xlim[1]) # Step 1: Assign group IDs to each continuous block of True/False in the mask group_ids = mask.ne(mask.shift()).cumsum() # Step 2: Find the group ID of the first valid continuous block first_valid_group = group_ids[mask].min() # Step 3: Filter only rows from the first valid block result = df[mask & (group_ids == first_valid_group)] print(result)
How This Works
Let's break down the key parts:
mask.ne(mask.shift()): This checks if the current row's mask value is different from the previous row. ATruehere marks the start of a new continuous block.cumsum(): We accumulate these boolean values to assign a unique ID to each continuous block (both valid and invalid).group_ids[mask].min(): We grab the smallest group ID from all valid rows, which corresponds to the first continuous block of rows in your target range.- Finally, we filter the DataFrame to keep only rows that are both in the target range and part of this first block.
Note on Your Expected Output
Running the code above will return rows with indices 2, 3, 4, and 5 (since x=6 falls within (2,6]). If you specifically want only indices 2-4, you likely intended the range to be (2,5] instead of (2,6]—just adjust xlim = [2,5] and the result will match your expected output exactly.
内容的提问来源于stack exchange,提问作者Leo
相关产品推荐
相关产品推荐

