如何用Pandas从CSV文件提取符合指定条件的唯一Line Source值?
使用Pandas实现CSV数据的筛选与唯一值提取
需求回顾
从包含空值的CSV文件中完成以下操作:
- 筛选
Name列值为'Mahmut'的记录 - 进一步筛选其中
Line Source的数值不在整个数据集的Line Destination列中的记录 - 提取这些记录里唯一的
Line Source值(预期输出:14)
示例数据
Id Name Library Line Source Line Destination 0 59 Ayla 2.0 57 34 1 60 Mahmut 2.0 14 22 2 61 Mine 2.0 22 43 3 62 Greg 2.0 14 62 4 63 Mahmut 2.0 14 33 5 64 Fiko 2.0 33 82 6 65 Jasmin 82 27 7 66 Mahmut 2.0 43 11 8 67 Ashley 2.0 62 53
实现步骤与代码
导入Pandas并读取CSV
直接读取CSV文件,Pandas会自动处理示例中的空值(比如第6行的Library空值),不影响目标字段的处理:import pandas as pd # 替换为你的CSV文件路径 df = pd.read_csv('your_file.csv')筛选Name为'Mahmut'的记录
用布尔索引快速定位目标用户的所有记录:mahmut_records = df[df['Name'] == 'Mahmut']筛选Line Source不在Line Destination列中的记录
先把整个数据集的Line Destination值转成集合(提升查询效率),再通过取反操作筛选符合条件的记录:# 排除Line Destination中的空值,避免干扰判断 dest_values = set(df['Line Destination'].dropna()) filtered_records = mahmut_records[~mahmut_records['Line Source'].isin(dest_values)]这里
~表示逻辑取反,isin()用于判断值是否在目标集合内。提取唯一的Line Source值
用unique()获取去重后的结果,再取出唯一值:result = filtered_records['Line Source'].unique()[0] print(result) # 输出:14
完整代码
import pandas as pd # 读取CSV文件 df = pd.read_csv('your_file.csv') # 筛选Name为Mahmut的记录 mahmut_df = df[df['Name'] == 'Mahmut'] # 获取所有Line Destination的非空值集合 dest_set = set(df['Line Destination'].dropna()) # 筛选Line Source不在目标集合中的记录 filtered = mahmut_df[~mahmut_df['Line Source'].isin(dest_set)] # 提取唯一值 unique_line_source = filtered['Line Source'].unique()[0] print(unique_line_source)
结果验证
根据示例数据:
- Mahmut的3条记录中,Line Source分别为14、14、43
- 整个数据集的Line Destination值包括:34、22、43、62、33、82、27、11、53
- 其中14不在该列表内,43存在于列表中,因此最终筛选出的唯一Line Source值为14,符合预期。
内容的提问来源于stack exchange,提问作者fatihbm
相关产品推荐
相关产品推荐

