如何用Pandas实现当列值大于阈值时删除后续指定行数?
Fixing Your Pandas Row Deletion Logic
Hey there! Let's work through this step by step to get your code working correctly for both your original and updated requirements.
First, Let's Address Your Initial Code Issue
You wanted to delete the 3 rows after any row where Pout > 3, but your original code has a few key problems:
- You're looping through the values of
Poutinstead of their indices, so you can't actually locate which row triggered the condition. df.drop(df[3:])deletes all rows starting at index 3, not the 3 rows after the matching row.- Pandas'
drop()returns a new DataFrame by default—you need to useinplace=Trueor reassign the result to save changes.
Corrected Code for Original Requirement (Delete 3 Rows After Pout > 3)
import pandas as pd df = pd.read_csv(file, sep=',', usecols=['Iin', 'Iout','Pout']) print("Original Pout:\n", df['Pout']) n_rows_to_delete = 3 # Get all indices where Pout > 3 condition_indices = df[df['Pout'] > 3].index.tolist() # Process indices in reverse to avoid index shifts messing up later deletions for idx in reversed(condition_indices): start = idx + 1 end = idx + n_rows_to_delete # Don't go beyond the last row of the DataFrame end = min(end, df.index[-1]) if start <= end: df.drop(range(start, end + 1), inplace=True) print("\nPout after deletion:\n", df['Pout'])
Updating for Your New Requirement (Delete 5 Rows After c > 3)
Your updated code sets values to None instead of deleting rows, and the break stops after the first matching row—let's fix that to handle all matching rows properly:
Optimized Code for Updated Requirement
import pandas as pd df = pd.DataFrame({ 'a': [1,5.0,1,2.3,2.1,2,1,3,4,7], 'b': [1,4,0.2,4.5,8.2,1,2,3,4,7], 'c': [1,4.5,5.4,6,2,4,2,3,4,7] }) n_rows_to_delete = 5 # Get all indices where column 'c' is greater than 3 condition_indices = df[df['c'] > 3].index.tolist() # Reverse the indices so deleting later rows doesn't affect earlier index positions for idx in reversed(condition_indices): start_delete = idx + 1 end_delete = idx + n_rows_to_delete # Ensure we don't try to delete rows that don't exist end_delete = min(end_delete, df.index[-1]) if start_delete <= end_delete: # Delete the range of rows df.drop(range(start_delete, end_delete + 1), inplace=True) print("Updated column 'c' after deletion:\n", df['c'])
Key Notes on This Approach
- Reverse Index Processing: By iterating over the condition indices in reverse order, we avoid issues where deleting rows shifts the indices of remaining rows—this ensures we target the correct rows every time.
- Boundary Check: The
min(end_delete, df.index[-1])makes sure we don't try to delete rows beyond the end of the DataFrame if there aren't 5 rows left after a matching entry. - Batch Deletion: Instead of modifying rows one by one (like your
atcalls), we delete full ranges at once, which is much more efficient for larger datasets.
内容的提问来源于stack exchange,提问作者Mera Louder
相关产品推荐
相关产品推荐

