如何从Pandas DataFrame列表列中移除对应值为NaN的客户编号?
问题分析与解决
你的代码存在两个核心问题:
- 逻辑反转:你需要移除的是
values列取NaN的客户编号,但当前nan_list获取的却是values非NaN的客户(values.notna()),完全搞反了筛选条件。 - 列表推导式错误:你直接判断整个
neighbors子列表是否在nan_list中,而没有遍历子列表里的每个客户编号进行过滤,导致无法移除目标元素。
修正后的代码
import pandas as pd import numpy as np def removing_nan_neighboors(custom_df): # 获取values为NaN的客户编号,这些是需要从neighbors中移除的对象 to_remove = custom_df[custom_df['values'].isna()]['customer'].tolist() # 遍历每个neighbors列表,过滤掉属于to_remove的客户编号 custom_df['neighbors'] = [ [neigh for neigh in neigh_list if neigh not in to_remove] for neigh_list in custom_df['neighbors'] ] return custom_df # 测试数据 customer = [1, 2, 3, 4, 5, 6] values = [np.nan, np.nan, 10, np.nan, 11, 12] neighbors = [[6, 2], [1, 3], [2, 4], [3, 5], [4, 6], [5, 1]] df = pd.DataFrame({'customer': customer, 'values': values, 'neighbors': neighbors}) df = removing_nan_neighboors(df) print(df)
输出结果
customer values neighbors 0 1 NaN [6] 1 2 NaN [3] 2 3 10.0 [] 3 4 NaN [3, 5] 4 5 11.0 [6] 5 6 12.0 [5]
完全符合你的预期输出。
内容的提问来源于stack exchange,提问作者enriicoo
相关产品推荐
相关产品推荐

