无需使用df.iterrows()遍历DataFrame且保留当前索引的替代方案
可行解决方案
有两种成熟方案可实现需求,均可以去掉外层iterrows循环,且保留DataFrame原始索引:
方案1:行级apply实现(改造成本最低,完全保留内层逻辑)
该方案运行效率比iterrows高5~10倍,且不需要修改你现有内层自定义判断逻辑,只需要把循环体封装为独立函数即可:
import pandas as pd attributes = ['attr1', 'attr2', 'attr3'] d = {'attr1': [1, 2], 'attr2': [3, 4], 'attr3' : [5, 6], 'meta': ['foo', 'bar']} df = pd.DataFrame(data=d) threshold = 0.5 weights = [0.3 , 0.3, 0.4] def check_row(row): results = [] # 完全保留原有的内层循环逻辑,无需任何修改 for attr in attributes: if attr == 'attr1': results.append(row[attr] * 5) else: results.append(row[attr]) confidence_level = sum(x * y for x, y in zip(results, weights)) / len(results) return confidence_level >= threshold # 直接获取符合条件的索引列表 indices = df[df.apply(check_row, axis=1)].index.tolist()
方案2:全向量化实现(性能最优,适合超大型DataFrame)
如果你的DataFrame规模在百万行以上,推荐用全向量化运算,性能比apply还要高几十到上百倍,只需要把内层的属性判断逻辑转换为列级运算即可:
import pandas as pd attributes = ['attr1', 'attr2', 'attr3'] d = {'attr1': [1, 2], 'attr2': [3, 4], 'attr3' : [5, 6], 'meta': ['foo', 'bar']} df = pd.DataFrame(data=d) threshold = 0.5 weights = [0.3 , 0.3, 0.4] # 对应原内层逻辑的属性转换 transformed_df = df[attributes].copy() transformed_df['attr1'] = transformed_df['attr1'] * 5 # 向量化计算置信度 confidence = (transformed_df * weights).sum(axis=1) / len(attributes) # 取符合条件的索引 indices = df[confidence >= threshold].index.tolist()
内容的提问来源于stack exchange,提问作者user17185996
相关产品推荐
相关产品推荐

