使用csv.DictReader读取TXT首列报KeyError: 'trip_id'的解决问询
读取CSV文件时包含trip_id列触发KeyError的问题排查与解决
问题背景
尝试读取TXT格式的CSV文件并提取指定列保存到新文件,只要columns_to_keep包含首列trip_id就触发KeyError: 'trip_id',移除该列后代码运行正常,且已确认列名拼写无误。
文件内容
trip_id,arrival_time,departure_time,stop_id,stop_sequence,stop_headsign,pickup_type,drop_off_type,shape_dist_traveled "2.T0.2-577-j23-1.2.R","07:29:00","07:29:00","at:46:1888:0:1","1","","0","0","0.00" "2.T0.2-577-j23-1.2.R","07:30:00","07:30:00","at:46:1885:0:1","2","","0","0","1732.64" "2.T0.2-577-j23-1.2.R","07:31:00","07:31:00","at:46:1884:0:1","3","","0","0","2589.62"
代码片段
def extract_and_save_columns(input_file, output_file, columns_to_keep): with open(input_file, "r", encoding='utf8') as source: reader = csv.DictReader(source) with open(output_file, "w", encoding='utf8', newline='') as result: fieldnames = columns_to_keep writer = csv.DictWriter(result, fieldnames=fieldnames) writer.writeheader() for row in reader: new_row = {col: row[col] for col in columns_to_keep} writer.writerow(new_row) # Keep columns of stop_times.txt stop_times_columns_to_keep = ["trip_id", "arrival_time", "departure_time", "stop_sequence", "shape_dist_traveled"] stop_times_input_file = path_to_gtfs + "stop_times.txt" stop_times_output_file = path_to_gtfs + "stop_times_filtered.txt" # Write new stop_times.txt extract_and_save_columns(stop_times_input_file, stop_times_output_file, stop_times_columns_to_keep)
报错信息
Traceback (most recent call last): File "...\PycharmProjects\GTFS_ST&T\main.py", line 48, in <module> extract_and_save_columns(stop_times_input_file, stop_times_output_file, stop_times_columns_to_keep) File "...\PycharmProjects\GTFS_ST&T\main.py", line 29, in extract_and_save_columns new_row = {col: row[col] for col in columns_to_keep} ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "...\PycharmProjects\GTFS_ST&T\main.py", line 29, in <dictcomp> new_row = {col: row[col] for col in columns_to_keep} ~~~^^^^^ KeyError: 'trip_id'
期望输出
trip_id,arrival_time,departure_time,stop_sequence,shape_dist_traveled "2.T0.2-577-j23-1.2.R","07:29:00","07:29:00","1","0.00" "2.T0.2-577-j23-1.2.R","07:30:00","07:30:00","2","1732.64" "2.T0.2-577-j23-1.2.R","07:31:00","07:31:00","3","2589.62"
排查原因
核心问题是CSV文件表头存在隐形字符:
- 最常见的是UTF-8 BOM(字节顺序标记,显示为
\ufeff),部分编辑器保存UTF-8文件时会自动添加; - 表头列名前后可能带有空格或不可见控制字符,导致
csv.DictReader解析出的列名与代码中指定的trip_id不匹配。
可以通过在代码中打印解析到的表头验证:
def extract_and_save_columns(input_file, output_file, columns_to_keep): with open(input_file, "r", encoding='utf8') as source: reader = csv.DictReader(source) print("解析到的表头:", reader.fieldnames) # 查看实际列名 # 后续代码不变
如果输出的trip_id前带有\ufeff或空格,即可确认问题。
解决方法
方案1:自动去除UTF-8 BOM(推荐)
将文件打开编码改为utf-8-sig,该编码会自动识别并去除UTF-8 BOM:
with open(input_file, "r", encoding='utf-8-sig') as source:
方案2:清洗表头隐形字符
如果是空格或其他控制字符,手动清洗每个列名:
reader = csv.DictReader(source) # 移除列名两端的所有空白和隐形控制字符 reader.fieldnames = [col.strip() for col in reader.fieldnames]
方案3:组合处理(覆盖所有场景)
结合两种方法,确保兼容BOM和其他隐形字符问题,修改后的完整代码:
def extract_and_save_columns(input_file, output_file, columns_to_keep): with open(input_file, "r", encoding='utf-8-sig') as source: reader = csv.DictReader(source) # 清洗表头,移除两端空白和隐形字符 cleaned_fieldnames = [col.strip() for col in reader.fieldnames] reader.fieldnames = cleaned_fieldnames with open(output_file, "w", encoding='utf8', newline='') as result: writer = csv.DictWriter(result, fieldnames=columns_to_keep) writer.writeheader() for row in reader: new_row = {col: row[col] for col in columns_to_keep} writer.writerow(new_row) # Keep columns of stop_times.txt stop_times_columns_to_keep = ["trip_id", "arrival_time", "departure_time", "stop_sequence", "shape_dist_traveled"] stop_times_input_file = path_to_gtfs + "stop_times.txt" stop_times_output_file = path_to_gtfs + "stop_times_filtered.txt" # Write new stop_times.txt extract_and_save_columns(stop_times_input_file, stop_times_output_file, stop_times_columns_to_keep)
内容的提问来源于stack exchange,提问作者Stenivic
相关产品推荐
相关产品推荐

