You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用pd.read_csv读取CSV时出现KeyError:列'TG'无法识别的解决方法

解决pandas读取CSV时存在'TG'列却触发KeyError的问题

问题描述

使用pd.read_csv读取CSV文件时触发KeyError: 'TG',但确认文件中确实存在名为'TG'的列,相关代码及报错如下:

import pandas as pd
import warnings 
warnings.filterwarnings('ignore')
stars = pd.read_csv('C:\\excel\\pulsar_stars_new.csv', on_bad_lines='skip')
stars_train = stars[((stars['TG'] == 0) & (stars['MIP'] >= 94.6640625) & (stars['MIP'] <= 95.2890625))]

报错信息:

KeyError                                  Traceback (most recent call last)
File C:\\anaconda\\lib\\site-packages\\pandas\\core\\indexes\\base.py:3802, in Index.get_loc(self, key, method, tolerance)
3801 try:
-> 3802     return self._engine.get_loc(casted_key)
3803 except KeyError as err:

File C:\\anaconda\\lib\\site-packages\\pandas\\_libs\\index.pyx:138, in pandas._libs.index.IndexEngine.get_loc()

File C:\\anaconda\\lib\\site-packages\\pandas\\_libs\\index.pyx:165, in pandas._libs.index.IndexEngine.get_loc()

File pandas\\_libs\\hashtable_class_helper.pxi:5745, in pandas._libs.hashtable.PyObjectHashTable.get_item()

File pandas\\_libs\\hashtable_class_helper.pxi:5753, in pandas._libs.hashtable.PyObjectHashTable.get_item()

KeyError: 'TG'

The above exception was the direct cause of the following exception:

KeyError                                  Traceback (most recent call last)
Cell In[30], line 1
----> 1 stars_train = stars[((stars['TG'] == 0) & (stars['MIP'] >= 94.6640625) & (stars['MIP'] <= 95.2890625))]

File C:\\anaconda\\lib\\site-packages\\pandas\\core\\frame.py:3807, in DataFrame.__getitem__(self, key)
3805 if self.columns.nlevels > 1:
3806     return self._getitem_multilevel(key)
-> 3807 indexer = self.columns.get_loc(key)
3808 if is_integer(indexer):
3809     indexer = [indexer]

File C:\\anaconda\\lib\\site-packages\\pandas\\core\\indexes\\base.py:3804, in Index.get_loc(self, key, method, tolerance)
3802     return self._engine.get_loc(casted_key)
3803 except KeyError as err:
-> 3804     raise KeyError(key) from err
3805 except TypeError:
3806     # If we have a listlike key, _check_indexing_error will raise
3807     #  InvalidIndexError. Otherwise we fall through and re-raise
3808     #  the TypeError.
3809     self._check_indexing_error(key)

KeyError: 'TG'

解决方案

1. 检查并清理列名中的空格/不可见字符

列名可能存在首尾空格、制表符或其他不可见字符,导致匹配失败:

  • 先查看当前所有列名:
    print(stars.columns.tolist())
    
  • 清理列名,去除首尾空格和特殊字符:
    # 去除列名首尾空格
    stars.columns = stars.columns.str.strip()
    
    # 若存在不可见字符,用正则清理
    import re
    stars.columns = [re.sub(r'[\s\t\r\n]+', '', col) for col in stars.columns]
    

2. 确认CSV的编码和分隔符

文件编码不正确或分隔符非默认逗号,会导致表头解析错误:

  • 指定正确的编码(如utf-8、gbk)和分隔符:
    stars = pd.read_csv('C:\\excel\\pulsar_stars_new.csv', on_bad_lines='skip', encoding='utf-8', sep=',')
    
  • 若不确定分隔符,可使用sep='\s+'匹配任意空白符,或启用delim_whitespace=True参数。

3. 检查on_bad_lines='skip'是否误跳过表头

该参数可能会跳过包含表头的行,导致DataFrame未正确读取表头:

  • 查看数据前几行,确认表头是否正确:
    print(stars.head())
    
  • 强制指定第一行为表头:
    stars = pd.read_csv('C:\\excel\\pulsar_stars_new.csv', on_bad_lines='skip', header=0)
    

4. 统一列名大小写

若文件中列名实际为小写(如'tg')或大小写混合,会导致匹配失败:

  • 统一列名大小写后再访问:
    # 转小写
    stars.columns = stars.columns.str.lower()
    stars_train = stars[((stars['tg'] == 0) & (stars['mip'] >= 94.6640625) & (stars['mip'] <= 95.2890625))]
    
    # 或转大写
    stars.columns = stars.columns.str.upper()
    

内容的提问来源于stack exchange,提问作者Karambombom

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 04:35:21