You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Pandas清洗CSV数据触发KeyError问题排查与解决

Pandas处理推特CSV数据时KeyError('tweet')问题排查与解决

问题场景

使用Pandas清洗CSV格式的推特数据,尝试将tweet列内容转为小写时触发KeyError错误,相关代码及报错信息如下:

原实现代码

import pandas as pd
import seaborn as sb
import matplotlib.pyplot as plt

# 导入数据集
df = pd.read_csv('tweet.csv')

# 生成缺失值热力图
plt.figure(figsize = (8,6))
sb.heatmap(df.isnull(), cbar=False , cmap = 'magma')

# 将所有推文转为小写
df['tweet'] = df['tweet'].apply(str.lower)
df['tweet'].head()

报错信息

KeyError                                  Traceback (most recent call last)
File ~\anaconda3\Lib\site-packages\pandas\core\indexes\base.py:3802, in Index.get_loc(self, key, method, tolerance)
   3801 try:
-> 3802     return self._engine.get_loc(casted_key)
   3803 except KeyError as err:

File ~\anaconda3\Lib\site-packages\pandas\_libs\index.pyx:138, in pandas._libs.index.IndexEngine.get_loc()

File ~\anaconda3\Lib\site-packages\pandas\_libs\index.pyx:165, in pandas._libs.index.IndexEngine.get_loc()

File pandas\_libs\hashtable_class_helper.pxi:5745, in pandas._libs.hashtable.PyObjectHashTable.get_item()

File pandas\_libs\hashtable_class_helper.pxi:5753, in pandas._libs.hashtable.PyObjectHashTable.get_item()

KeyError: 'tweet'

The above exception was the direct cause of the following exception:

KeyError                                  Traceback (most recent call last)
Cell In[2], line 13
     10 sb.heatmap(df.isnull(), cbar=False , cmap = 'magma')
     12 #convert all tweet into lowercase
---> 13 df['tweet'] = df['tweet'].apply(str.lower)
     14 df['tweet'].head()

File ~\anaconda3\Lib\site-packages\pandas\core\frame.py:3807, in DataFrame.__getitem__(self, key)
   3805 if self.columns.nlevels > 1:
   3806     return self._getitem_multilevel(key)
-> 3807 indexer = self.columns.get_loc(key)
   3808 if is_integer(indexer):
   3809     indexer = [indexer]

File ~\anaconda3\Lib\site-packages\pandas\core\indexes\base.py:3804, in Index.get_loc(self, key, method, tolerance)
   3802     return self._engine.get_loc(casted_key)
   3803 except KeyError as err:
-> 3804     raise KeyError(key) from err
   3805 except TypeError:
   3806     # If we have a listlike key, _check_indexing_error will raise
   3807     #  InvalidIndexError. Otherwise we fall through and re-raise
   3808     #  the TypeError.
   3809     self._check_indexing_error(key)

KeyError: 'tweet'

问题原因

KeyError: 'tweet'明确表示当前DataFrame中不存在名为'tweet'的列,常见触发原因包括:

  • CSV文件中的列名拼写与代码中使用的'tweet'不一致(比如大小写差异:'Tweet'、'TWEET',或者包含空格:' tweet ')
  • CSV文件使用了非逗号的分隔符(比如制表符\t),导致Pandas读取时未正确识别列
  • CSV文件的表头行缺失,导致列名被识别为第一行数据

解决方案与修正代码

步骤1:排查实际列名

在读取CSV后,先打印当前DataFrame的列名,确认实际列名:

print(df.columns.tolist())

步骤2:针对不同情况处理

情况1:列名存在大小写/空格差异

如果输出的列名是'Tweet'或者' tweet '这类,可以统一标准化列名:

# 去除列名前后空格并转为小写
df.columns = df.columns.str.strip().str.lower()

情况2:CSV分隔符错误

如果CSV用的是制表符或其他分隔符,读取时指定sep参数:

# 比如制表符分隔的CSV
df = pd.read_csv('tweet.csv', sep='\t')

情况3:表头缺失

如果CSV没有表头行,读取时指定header=None并手动设置列名:

df = pd.read_csv('tweet.csv', header=None)
# 假设推文列是第2列(索引1),根据实际数据补充列名
df.columns = ['id', 'tweet', 'created_at']

修正后的完整代码

import pandas as pd
import seaborn as sb
import matplotlib.pyplot as plt

# 导入数据集,先处理列名问题
df = pd.read_csv('tweet.csv')
# 标准化列名:去空格+转小写
df.columns = df.columns.str.strip().str.lower()

# 生成缺失值热力图
plt.figure(figsize = (8,6))
sb.heatmap(df.isnull(), cbar=False , cmap = 'magma')

# 将所有推文转为小写
df['tweet'] = df['tweet'].apply(str.lower)
print(df['tweet'].head())

内容的提问来源于stack exchange,提问作者Nur Hidayah Athira

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 14:17:50