如何在遍历pandas DataFrame时动态调整起始行、跳过指定行
问题场景
现有存储1-10数值的pandas DataFrame,需要逐行遍历数据,匹配到数值5时跳过下一行后恢复遍历。
原有测试代码如下:
import pandas as pd df=pd.DataFrame([1,2,3,4,5,6,7,8,9,10]) df.columns=['number'] start=0 for index, row in df.iloc[start:].iterrows(): print(index, row['number']) if row['number']==5: start=index+2
原有代码实际输出:
0 1 1 2 2 3 3 4 4 5 5 6 6 7 7 8 8 9 9 10
期望输出:
0 1 1 2 2 3 3 4 4 5 6 7 7 8 8 9 9 10
失效原因
Python的for循环在启动前会完成in关键字后可迭代对象的求值,生成固定的迭代序列。原有代码中df.iloc[start:]在循环启动时就已经按初始值start=0生成了全量行的迭代器,后续在循环内修改start变量不会改变已经生成的迭代序列,因此跳转逻辑不生效。
正确实现方案
方案1:手动控制索引的while循环(灵活性最高,支持任意行数前后跳转)
import pandas as pd df = pd.DataFrame([1,2,3,4,5,6,7,8,9,10]) df.columns = ['number'] index = 0 row_count = len(df) while index < row_count: current_row = df.iloc[index] print(index, current_row['number']) if current_row['number'] == 5: # 匹配到目标值,直接跳过下一行,索引+2 index += 2 else: # 未匹配到目标值,正常遍历索引+1 index += 1
方案2:跳过标记配合iterrows遍历(适合仅需跳过后续固定行数的简单场景)
import pandas as pd df = pd.DataFrame([1,2,3,4,5,6,7,8,9,10]) df.columns = ['number'] # 标记是否需要跳过当前行 skip_current = False for index, row in df.iterrows(): if skip_current: skip_current = False continue print(index, row['number']) if row['number'] == 5: # 标记下一行需要跳过 skip_current = True
两种方案运行后均可得到期望输出。
内容的提问来源于stack exchange,提问作者new coding
相关产品推荐
相关产品推荐

