如何用Python移除CSV文件注释行以实现Pandas正常解析
解决Pandas读取带#注释CSV文件的ParserError问题
问题情况
执行以下代码读取SeaLevel.csv时触发ParserError:
import pandas as pd sea_level_df = pd.read_csv(r"C:\Users\slaye\OneDrive\Desktop\SeaLevel.csv") display(sea_level_df)
该CSV文件前3行是#开头的注释行,文件内容如下:
#title = 全球海洋(南纬66度至北纬66度)平均海平面异常(保留年度信号) #institution = NOAA/卫星测高实验室 #references = NOAA海平面上升数据 year,TOPEX/Poseidon,Jason-1,Jason-2,Jason-3 1992.9614,-16.27000, 1992.9865,-17.97000, 1993.0123,-14.87000, 1993.0407,-19.87000, 1993.0660,-25.27000, 1993.0974,-29.37000,
报错信息(翻译后):
ParserError 回溯(最近的调用最后) 输入 In [14], 在<cell line: 2>()代码块中 1 import pandas as pd ----> 2 sea_level_df = pd.read_csv(r"C:\Users\slaye\OneDrive\Desktop\SeaLevel.csv") 3 display(sea_level_df) 文件 ~\anaconda3\lib\site-packages\pandas\util\_decorators.py:311, 在deprecate_nonkeyword_arguments.<locals>.decorate.<locals>.wrapper(*args, **kwargs)函数中 305 如果参数数量超过允许的数量: 306 发出警告: 307 msg.format(arguments=arguments), 308 FutureWarning, 309 stacklevel=stacklevel, 310 ) --> 311 返回原函数执行结果 文件 ~\anaconda3\lib\site-packages\pandas\io\parsers\readers.py:680, 在read_csv函数中 665 kwds_defaults = _refine_defaults_read( 666 dialect, 667 delimiter, (...) 676 defaults={"delimiter": ","}, 677 ) 678 更新kwds参数 --> 680 返回_read(filepath_or_buffer, kwds) 文件 ~\anaconda3\lib\site-packages\pandas\io\parsers\readers.py:581, 在_read函数中 578 返回parser对象 580 使用parser上下文: --> 581 返回parser.read(nrows) 文件 ~\anaconda3\lib\site-packages\pandas\io\parsers\readers.py:1254, 在TextFileReader.read(self, nrows)方法中 1252 验证nrows为整数 1253 尝试执行: -> 1254 index, columns, col_dict = self._engine.read(nrows) 1255 捕获异常: 1256 关闭parser 文件 ~\anaconda3\lib\site-packages\pandas\io\parsers\c_parser_wrapper.py:225, 在CParserWrapper.read(self, nrows)方法中 223 尝试执行: 224 如果启用低内存模式: --> 225 chunks = self._reader.read_low_memory(nrows) 226 # 合并chunks数据 227 data = _concatenate_chunks(chunks) 文件 ~\anaconda3\lib\site-packages\pandas\_libs\parsers.pyx:805, 在pandas._libs.parsers.TextReader.read_low_memory()中 文件 ~\anaconda3\lib\site-packages\pandas\_libs\parsers.pyx:861, 在pandas._libs.parsers.TextReader._read_rows()中 文件 ~\anaconda3\lib\site-packages\pandas\_libs\parsers.pyx:847, 在pandas._libs.parsers.TextReader._tokenize_rows()中 文件 ~\anaconda3\lib\site-packages\pandas\_libs\parsers.pyx:1960, 在pandas._libs.parsers.raise_parser_error()中 ParserError: 数据分词错误。C层错误:第4行预期1个字段,实际检测到5个
解决方案
无需手动修改CSV文件,直接通过pandas.read_csv()的参数即可跳过注释行,两种实用方法:
方法一:使用comment参数
指定#为注释标记,Pandas会自动忽略所有以#开头的行:
import pandas as pd sea_level_df = pd.read_csv(r"C:\Users\slaye\OneDrive\Desktop\SeaLevel.csv", comment='#') display(sea_level_df)
方法二:使用skiprows参数
明确指定跳过前3行注释内容:
import pandas as pd sea_level_df = pd.read_csv(r"C:\Users\slaye\OneDrive\Desktop\SeaLevel.csv", skiprows=3) display(sea_level_df)
推荐使用comment参数,它更灵活——即使后续文件注释行数量变化,也无需修改代码。
内容的提问来源于stack exchange,提问作者Bijan Khair Rahman
相关产品推荐
相关产品推荐

