如何使用Pandas替代原生csv库筛选CSV指定行并导出带自定义表头的文件
原生CSV处理逻辑转Pandas实现需求
原有功能说明
现有基于Python原生csv库的代码,实现以下功能:
- 读取原始CSV文件,筛选前三个字段为
Data、9、record的行 - 导出结果的表头依次为
timestamp、latitude、longitude、distance、altitude、speed - 字段取值来自原始数据的第5、8、11、14、17、20列(列索引从1开始计数),同时去除字段两端的双引号
- 最后读取导出的文件打印每行内容验证结果
原生csv库实现代码
import csv with open("assets/ride.csv", "r") as source: lines = source.readlines() with open("solution06.csv", "w") as new_file: # 写入表头 new_file.write(','.join(('timestamp','latitude','longitude','distance','altitude','speed'))) new_file.write('\n') # 遍历原始表头之后的所有行 for line in lines[1:]: line = line.split(',') # 仅保留以"Data,<number>,record"开头的行 if line[0] == 'Data'and line[1] == '9' and line[2] == 'record': # 拼接需要的字段并写入新行 new_file.write(','.join(line[column - 1].strip('"') for column in (5, 8, 11, 14, 17, 20))) new_file.write('\n') with open("solution06.csv", "r") as source: for line in source.readlines(): print(line.strip('\n').split(','))
等效Pandas实现代码
import pandas as pd # 读取原始CSV文件,不自动识别表头,所有字段以字符串格式读取避免格式转换异常 df = pd.read_csv("assets/ride.csv", header=None, dtype=str) # 筛选符合条件的行:前3列依次为Data、9、record,同时跳过原始第一行(对应原代码lines[1:]) filtered_df = df.loc[1:][(df[0] == 'Data') & (df[1] == '9') & (df[2] == 'record')] # 取指定列(原1起始索引5/8/11/14/17/20对应Pandas0起始索引4/7/10/13/16/19),去除所有字段两端双引号,重命名表头 result_df = filtered_df[[4,7,10,13,16,19]].apply(lambda x: x.str.strip('"')) result_df.columns = ['timestamp','latitude','longitude','distance','altitude','speed'] # 导出为CSV文件,不保存索引 result_df.to_csv("solution06.csv", index=False) # 读取导出文件打印验证 verify_df = pd.read_csv("solution06.csv", dtype=str) for _, row in verify_df.iterrows(): print(row.tolist())
内容的提问来源于stack exchange,提问作者user12755836
相关产品推荐
相关产品推荐

