Python批量处理CSV文件:移除第二列功能实现求助
问题:移除Event CSV文件的第二列数据
需要为现有Python程序添加功能,移除每个Event CSV文件的第二列(即Event #列)。尝试过相关方案但未成功,现有CSV文件格式如下:
Time/Date,Event #,Event Desc 05/19/2020 20:12:30,29,Advance Drive ON 05/19/2020 20:32:23,29,Advance Drive ON 05/19/2020 20:35:13,29,Advance Drive ON 05/19/2020 20:39:50,37,Discharge 1 Plug Chute Fault 05/19/2020 20:47:40,68,LMI is in OFF Mode
现有处理函数代码:
# A function to clean the Event Files of raw data def CleanEventFiles(EF_files, eventHeader, EFmachineID): logging.debug(f'Cleaning Event files...') # Write to program logger for f in EF_files: # FOR ALL FILES IN EVENT FILES IsFileReadOnly(f) # check to see if the file is READ ONLY print(f'\nCleaning file: {f}') # tell user which file is being cleaned print('\tReplacing new MachineIDs & File Headers...') # print stuff to the user logging.debug(f'\tReplacing headers for file {f}') # write to program logger with open(f, newline='', encoding='latin-1') as g: # open file as read r = csv.reader((line.replace('\0', '') for line in g)) # declare read variable while removing NULLs next(r) # remove old machineID data = [line for line in r] # set list to all data in file data[0] = eventHeader # replace first line with new header data.insert(0, EFmachineID) # add line before header for machine ID WriteData(f, data) # write data to the file
尝试过添加del r[1]之类的代码,但仅能移除Event #表头,对应的数据列仍保留,求正确移除第二列的方法。
解决方案
问题核心是csv.reader返回的是迭代器而非列表,直接操作r无法批量处理所有行。要彻底移除第二列,需在读取每一行时剔除索引为1的元素(Python索引从0开始)。
修改后的完整代码:
# A function to clean the Event Files of raw data def CleanEventFiles(EF_files, eventHeader, EFmachineID): logging.debug(f'Cleaning Event files...') # Write to program logger for f in EF_files: # FOR ALL FILES IN EVENT FILES IsFileReadOnly(f) # check to see if the file is READ ONLY print(f'\nCleaning file: {f}') # tell user which file is being cleaned print('\tReplacing new MachineIDs & File Headers...') # 修正转义字符问题 print('\tRemoving second column (Event #)...') # 新增操作提示 logging.debug(f'\tReplacing headers and removing second column for file {f}') # 更新日志信息 with open(f, newline='', encoding='latin-1') as g: # open file as read r = csv.reader((line.replace('\0', '') for line in g)) # declare read variable while removing NULLs next(r) # remove old machineID # 遍历每一行,保留第0列和第2列及以后的内容,剔除第二列 data = [row[:1] + row[2:] for row in r] data[0] = eventHeader # replace first line with new header data.insert(0, EFmachineID) # add line before header for machine ID WriteData(f, data) # write data to the file
关键修改说明:
- 替换数据读取逻辑:将
data = [line for line in r]改为data = [row[:1] + row[2:] for row in r],对每一行进行切片处理,直接剔除第二列数据。 - 修正原代码中的
&为&,确保控制台提示正常显示。 - 新增控制台提示和日志内容,让操作过程更清晰。
注意:需确保传入的eventHeader是移除第二列后的表头,例如['Time/Date', 'Event Desc'],否则会出现表头与数据列数不匹配的问题。
内容的提问来源于stack exchange,提问作者Kyle Lucas
相关产品推荐
相关产品推荐

