如何不使用Pandas库删除CSV指定列?代码异常求助
不使用Pandas删除CSV指定列的代码修正
需求说明
不使用Pandas库,删除CSV文件中名为Department和Allocation的列,且无法保证这两列的固定位置。
示例CSV文件
Name,Age,YearofService,Department,Allocation Birla,49,12,Welding,Production Robin,38,10,Molding,Production
错误代码
with open('input.csv','r') as i: with open('output.csv','w',newline='') as o: reader=csv.reader(i) writer = csv.writer(o) for row in reader: for i in range(len(row)): if row[i]!="Department" and row[i]!="Allocation": writer.writerow(row)
当前错误输出
Name Birla Robin Age 49 38 YearofService 12 10
期望正确输出
Name,Age,YearofService Birla,49,12 Robin,38,10
问题分析
原代码逻辑错误:遍历列时,只要当前列不是目标列就把整行写入,导致同一行被重复写入多次,且没有过滤掉目标列的内容。正确思路应该是先通过表头确定需要保留的列索引,后续每一行仅保留这些索引对应的内容。
修正后的代码
import csv # 定义需要删除的列名集合 columns_to_remove = {"Department", "Allocation"} with open('input.csv', 'r') as infile, open('output.csv', 'w', newline='') as outfile: reader = csv.reader(infile) writer = csv.writer(outfile) # 读取表头,筛选出需要保留的列索引 header = next(reader) keep_indices = [idx for idx, col_name in enumerate(header) if col_name not in columns_to_remove] # 写入过滤后的表头 writer.writerow([header[idx] for idx in keep_indices]) # 遍历处理每一行数据,仅保留指定索引的列 for row in reader: filtered_row = [row[idx] for idx in keep_indices] writer.writerow(filtered_row)
代码说明
- 先读取CSV的表头行,通过列表推导式筛选出不需要删除的列的索引,存入
keep_indices - 先将过滤后的表头写入输出文件
- 遍历后续每一行数据,根据
keep_indices提取对应位置的元素,组成新行后写入输出文件
这种方法无需依赖列的固定位置,能准确删除指定列,且不会出现重复写入的问题。
内容的提问来源于stack exchange,提问作者Balaji R B
相关产品推荐
相关产品推荐

