JupyterLab读取CSV触发IndexError:列表索引越界求助
问题描述
将Excel保存为CSV格式后,在JupyterLab中用pandas读取时触发IndexError: list index out of range,所用代码和同学完全一致,但仅自己出现该错误。
所用代码
Use_raw = pd.read_csv("U15_US_2021.csv", header=3, index_col=1, na_values='---').drop('Unnamed: 0', axis=1) Use=Use_raw.iloc[0:17,0:15].fillna(0) ValueAdded=Use_raw.iloc[18:22,0:15].fillna(0) FinalDemand=Use_raw.iloc[0:17,16:21].fillna(0) Supply_raw = pd.read_csv('S15_US_2021.csv', header=3, index_col=1,na_values='---') Supply=Supply_raw.iloc[0:17,1:16].fillna(0)
报错堆栈信息
IndexError Traceback (most recent call last) Cell In[23], line 1 ----> 1 Use_raw = pd.read_csv("U15_US_2021.csv", header=3, index_col=1, na_values='--- ').drop('Unnamed: 0', axis=1) 2 Use=Use_raw.iloc[0:17,0:15].fillna(0) 3 ValueAdded=Use_raw.iloc[18:22,0:15].fillna(0) File ~\anaconda3\envs\IO\lib\site-packages\pandas\util\_decorators.py:211, in deprecate_kwarg. <locals>._deprecate_kwarg.<locals>.wrapper(*args, **kwargs) 209 else: 210 kwargs[new_arg_name] = new_arg_value --> 211 return func(*args, **kwargs) File ~\anaconda3\envs\IO\lib\site-packages\pandas\util\_decorators.py:331, in deprecate_nonkeyword_arguments.<locals>.decorate.<locals>.wrapper(*args, **kwargs) 325 if len(args) > num_allow_args: 326 warnings.warn( 327 msg.format(arguments=_format_argument_list(allow_args)), 328 FutureWarning, 329 stacklevel=find_stack_level(), 330 ) --> 331 return func(*args, **kwargs) File ~\anaconda3\envs\IO\lib\site-packages\pandas\io\parsers\readers.py:950, in read_csv(filepath_or_buffer, sep, delimiter, header, names, index_col, usecols, squeeze, prefix, mangle_dupe_cols, dtype, engine, converters, true_values, false_values, skipinitialspace, skiprows, skipfooter, nrows, na_values, keep_default_na, na_filter, verbose, skip_blank_lines, parse_dates, infer_datetime_format, keep_date_col, date_parser, dayfirst, cache_dates, iterator, chunksize, compression, thousands, decimal, lineterminator, quotechar, quoting, doublequote, escapechar, comment, encoding, encoding_errors, dialect, error_bad_lines, warn_bad_lines, on_bad_lines, delim_whitespace, low_memory, memory_map, float_precision, storage_options) 935 kwds_defaults = _refine_defaults_read( 936 dialect, 937 delimiter, (...) 946 defaults={"delimiter": ","}, 947 ) 948 kwds.update(kwds_defaults) --> 950 return _read(filepath_or_buffer, kwds) File ~\anaconda3\envs\IO\lib\site-packages\pandas\io\parsers\readers.py:605, in _read(filepath_or_buffer, kwds) 602 _validate_names(kwds.get("names", None)) 604 # Create the parser. --> 605 parser = TextFileReader(filepath_or_buffer, **kwds) 607 if chunksize or iterator: 608 return parser File ~\anaconda3\envs\IO\lib\site-packages\pandas\io\parsers\readers.py:1442, in TextFileReader.__init__(self, f, engine, **kwds) 1439 self.options["has_index_names"] = kwds["has_index_names"] 1441 self.handles: IOHandles | None = None --> 1442 self._engine = self._make_engine(f, self.engine) File ~\anaconda3\envs\IO\lib\site-packages\pandas\io\parsers\readers.py:1753, in TextFileReader._make_engine(self, f, engine) 1750 raise ValueError(msg) 1752 try: --> 1753 return mapping[engine](f, **self.options) 1754 except Exception: 1755 if self.handles is not None: File ~\anaconda3\envs\IO\lib\site-packages\pandas\io\parsers\c_parser_wrapper.py:174, in CParserWrapper.__init__(self, src, **kwds) 164 if self._reader.leading_cols == 0 and is_index_col( 165 self.index_col # type: ignore[has-type] 166 ): 168 self._name_processed = True 169 ( 170 index_names, 171 # error: Cannot determine type of 'names' 172 self.names, # type: ignore[has-type] 173 self.index_col, --> 174 ) = self._clean_index_names( 175 # error: Cannot determine type of 'names' 176 self.names, # type: ignore[has-type] 177 # error: Cannot determine type of 'index_col' 178 self.index_col, # type: ignore[has-type] 179 ) 181 if self.index_names is None: 182 self.index_names = index_names File ~\anaconda3\envs\IO\lib\site-packages\pandas\io\parsers\base_parser.py:999, in ParserBase._clean_index_names(self, columns, index_col) 997 break 998 else: --> 999 name = cp_cols[c] 1000 columns.remove(name) 1001 index_names.append(name) IndexError: list index out of range
分析与解决
错误核心原因
这个IndexError是因为代码指定了index_col=1(将第2列作为索引),但pandas解析CSV时,发现第4行(header=3对应索引从0开始的第3行,即Excel里的第4行)的列数不足,找不到索引为1的列,因此触发报错。
和同学代码一致但报错,大概率是Excel转CSV的格式差异,而非代码本身问题,具体排查修复步骤如下:
排查修复步骤
检查CSV文件实际结构
用记事本打开U15_US_2021.csv,查看第4行的列数、分隔符,确认是否存在列缺失,以及分隔符是否为逗号(同学的文件应为逗号分隔)。临时调整代码确认数据结构
先去掉header、index_col等参数,读取整个文件查看实际内容:import pandas as pd df = pd.read_csv("U15_US_2021.csv") print(df.head(15)) # 打印前15行,确认表头和数据位置根据输出结果,再确定正确的
header行数和index_col列索引。修正Excel转CSV的保存设置
- 保存Excel时,选择**「CSV(逗号分隔)(*.csv)」**格式,不要选「CSV(制表符分隔)」或其他变体;
- 保存时确认编码为UTF-8(和同学保持一致,避免解析乱码导致列识别错误);
- 确保Excel文件中没有隐藏列,保存时所有可见列都被导出。
统一参数细节
报错信息里的na_values是'--- '(带空格),但代码里写的是'---',统一成一致的参数值,避免缺失值识别异常。指定分隔符(若需要)
如果你的CSV是制表符或分号分隔,在read_csv中添加sep参数,比如:# 制表符分隔 Use_raw = pd.read_csv("U15_US_2021.csv", header=3, index_col=1, na_values='---', sep='\t').drop('Unnamed: 0', axis=1) # 分号分隔 Use_raw = pd.read_csv("U15_US_2021.csv", header=3, index_col=1, na_values='---', sep=';').drop('Unnamed: 0', axis=1)
内容的提问来源于stack exchange,提问作者Anna Voldseth
相关产品推荐
相关产品推荐

