使用Python的csv.DictReader读取CSV文件时丢失列的问题
问题描述
我有如下格式的.csv文件,希望完整读取它:
--------------------------------------------------------------- #INFO | | | | | | --------------------------------------------------------------- #Study name | g1 | | | | | --------------------------------------------------------------- #Respondent Name | name | | | | | --------------------------------------------------------------- #Respondent Age | 20 | | | | | --------------------------------------------------------------- #DATA | | | | | | --------------------------------------------------------------- Row | Timestamp | Source | Event | Sample | Anger| --------------------------------------------------------------- 1 | 133 | Face | 1 | 3 | 0.44 | --------------------------------------------------------------- 2 | 240 | Face | 1 | 4 | 0.20 | --------------------------------------------------------------- 3 | 12 | Face | 1 | 5 | 0.13 | --------------------------------------------------------------- 4 | 133 | Face | 1 | 6 | 0.75 | --------------------------------------------------------------- 5 | 87 | Face | 1 | 7 | 0.25 | ---------------------------------------------------------------
我用以下Python代码读取文件并打印:
import csv def read_csv(): with open("in/2.csv", encoding="utf-8-sig") as f: reader = csv.DictReader(f) print(list(reader)) read_csv()
但输出只包含前两列:
[{'#INFO': '#Study name', '': ''}, {'#INFO': '#Respondent Name', '': ''}, {'#INFO': '#Respondent Age', '': ''}, {'#INFO': '#DATA', '': ''}, {'#INFO': 'Row', '': 'Anger'}, {'#INFO': '1', '': '4.40E-01'}, {'#INFO': '2', '': '2.00E-01'}, {'#INFO': '3', '': '1.30E-01'}, {'#INFO': '4', '': '7.50E-01'}, {'#INFO': '5', '': '2.50E-01'}]
为什么其余列会丢失?我想读取从Row开始的所有列,求指导。
问题原因与解决方法
原因
Python的csv.DictReader默认用逗号,作为分隔符,但你的文件是用|作为分隔符,且分隔符前后还有空格,默认解析逻辑只能识别前两列,后续列被错误合并或判定为空。
解决步骤
1. 指定正确的分隔符与空格处理
修改csv.DictReader参数,指定分隔符为|,并开启skipinitialspace=True忽略分隔符后的空格:
import csv def read_csv(): with open("in/2.csv", encoding="utf-8-sig") as f: reader = csv.DictReader(f, delimiter='|', skipinitialspace=True) print(list(reader)) read_csv()
2. 跳过无效行,只读取目标数据
上述代码会读取所有行,而你需要的是从Row开头的行开始的真实数据,因此需要先定位表头行,跳过前面的配置行和分隔线:
import csv def read_csv(): with open("in/2.csv", encoding="utf-8-sig") as f: lines = f.readlines() # 找到表头行(以Row开头的行) header_index = None for i, line in enumerate(lines): if line.strip().startswith('Row'): header_index = i break if header_index is not None: # 从表头行开始解析 reader = csv.DictReader(lines[header_index:], delimiter='|', skipinitialspace=True) data = [] for row in reader: # 过滤掉全是分隔符的空行 if not all(v == '' for v in row.values()): data.append(row) print(data) read_csv()
3. 最终效果
执行后会得到包含所有列的正确数据:
[{'Row': '1', 'Timestamp': '133', 'Source': 'Face', 'Event': '1', 'Sample': '3', 'Anger': '0.44'}, {'Row': '2', 'Timestamp': '240', 'Source': 'Face', 'Event': '1', 'Sample': '4', 'Anger': '0.20'}, {'Row': '3', 'Timestamp': '12', 'Source': 'Face', 'Event': '1', 'Sample': '5', 'Anger': '0.13'}, {'Row': '4', 'Timestamp': '133', 'Source': 'Face', 'Event': '1', 'Sample': '6', 'Anger': '0.75'}, {'Row': '5', 'Timestamp': '87', 'Source': 'Face', 'Event': '1', 'Sample': '7', 'Anger': '0.25'}]
内容的提问来源于stack exchange,提问作者ChenBr
相关产品推荐
相关产品推荐

