如何解决Pandas中read_csv出现的SystemError问题?
使用pd.read_csv()触发SystemError的排查与解决思路
问题场景
运行以下代码时触发底层SystemError:
import pandas as pd df = pd.read_csv('MLproject/color_names.csv', usecols=['Name', 'Hex'])
完整报错信息:
SystemError Traceback (most recent call last) Cell In[18], line 2 1 import pandas as pd ----> 2 df = pd.read_csv('MLproject/color_names.csv', usecols=['Name', 'Hex']) 3 print(df) File ~\AppData\Roaming\Python\Python312\site-packages\pandas\io\parsers\readers.py:1026, in read_csv(filepath_or_buffer, sep, delimiter, header, names, index_col, usecols, dtype, engine, converters, true_values, false_values, skipinitialspace, skiprows, skipfooter, nrows, na_values, keep_default_na, na_filter, verbose, skip_blank_lines, parse_dates, infer_datetime_format, keep_date_col, date_parser, date_format, dayfirst, cache_dates, iterator, chunksize, compression, thousands, decimal, lineterminator, quotechar, quoting, doublequote, escapechar, comment, encoding, encoding_errors, dialect, on_bad_lines, delim_whitespace, low_memory, memory_map, float_precision, storage_options, dtype_backend) 1013 kwds_defaults = _refine_defaults_read( 1014 dialect, 1015 delimiter, (...) 1022 dtype_backend=dtype_backend, 1023 ) 1024 kwds.update(kwds_defaults) -> 1026 return _read(filepath_or_buffer, kwds) File ~\AppData\Roaming\Python\Python312\site-packages\pandas\io\parsers\readers.py:620, in _read(filepath_or_buffer, kwds) 617 _validate_names(kwds.get("names", None)) 619 # Create the parser. -> 620 parser = TextFileReader(filepath_or_buffer, **kwds) 622 if chunksize or iterator: 623 return parser File ~\AppData\Roaming\Python\Python312\site-packages\pandas\io\parsers\readers.py:1620, in TextFileReader.__init__(self, f, engine, **kwds) 1617 self.options["has_index_names"] = kwds["has_index_names"] 1619 self.handles: IOHandles | None = None -> 1620 self._engine = self._make_engine(f, self.engine) File ~\AppData\Roaming\Python\Python312\site-packages\pandas\io\parsers\readers.py:1898, in TextFileReader._make_engine(self, f, engine) 1895 raise ValueError(msg) 1897 try: -> 1898 return mapping[engine](f, **self.options) 1899 except Exception: 1900 if self.handles is not None: File ~\AppData\Roaming\Python\Python312\site-packages\pandas\io\parsers\c_parser_wrapper.py:61, in CParserWrapper.__init__(self, src, **kwds) 60 def __init__(self, src: ReadCsvBuffer[str], **kwds) -> None: ---> 61 super().__init__(kwds) 62 self.kwds = kwds 63 kwds = kwds.copy() File ~\AppData\Roaming\Python\Python312\site-packages\pandas\io\parsers\base_parser.py:185, in ParserBase.__init__(self, kwds) 181 self._name_processed = False 183 self._first_chunk = True -> 185 self.usecols, self.usecols_dtype = self._validate_usecols_arg(kwds["usecols"]) 187 # Fallback to error to pass a sketchy test(test_override_set_noconvert_columns) 188 # Normally, this arg would get pre-processed earlier on 189 self.on_bad_lines = kwds.get("on_bad_lines", self.BadLineHandleMethod.ERROR) File ~\AppData\Roaming\Python\Python312\site-packages\pandas\io\parsers\base_parser.py:1026, in ParserBase._validate_usecols_arg(self, usecols) 1020 if not is_list_like(usecols): 1021 # see gh-20529 1022 # 1023 # Ensure it is iterable container but not string. 1024 raise ValueError(msg) -> 1026 usecols_dtype = lib.infer_dtype(usecols, skipna=False) 1028 if usecols_dtype not in ("empty", "integer", "string"): 1029 raise ValueError(msg) File lib.pyx:1628, in pandas._libs.lib.infer_dtype() SystemError: ../numpy/_core/src/multiarray/iterators.c:192: bad argument to internal function
30分钟前代码可正常运行,未做任何改动,已尝试重装Pandas/Numpy、安装Numpy 1.19.3、将usecols改为[0,1],均无效。
排查与解决步骤
重建虚拟环境
这类底层C扩展报错大概率是Python环境损坏导致,创建全新虚拟环境隔离问题:# Windows环境 python -m venv new_ml_env new_ml_env\Scripts\activate pip install pandas numpy在新环境中运行代码,验证是否恢复正常。
检查文件完整性与路径
虽然报错指向Numpy,但也可能是CSV文件损坏或路径解析问题:- 重新拷贝一份
color_names.csv到项目目录,确认文件未被篡改 - 使用绝对路径替代相对路径,避免路径解析异常:
import os file_path = os.path.abspath('MLproject/color_names.csv') df = pd.read_csv(file_path, usecols=['Name', 'Hex'])
- 重新拷贝一份
切换解析引擎
改用Python引擎绕过C扩展层问题:df = pd.read_csv('MLproject/color_names.csv', usecols=['Name', 'Hex'], engine='python')安装Python3.12兼容的版本组合
Python3.12较新,旧版本Numpy/Pandas兼容性较差,安装适配版本:# 安装适配Python3.12的稳定版本 pip install pandas==2.1.4 numpy==1.26.3清理残留文件后重装
卸载库后手动清理残留文件,避免缓存干扰:- 删除
~\AppData\Roaming\Python\Python312\site-packages下的pandas、numpy相关文件夹 - 清理pip缓存:
pip cache purge
之后重新安装指定版本的库。
- 删除
内容的提问来源于stack exchange,提问作者user372087
相关产品推荐
相关产品推荐

