Coursera数据分析作业Autograder报KeyError: 'STNAME'求解决
我在Coursera数据分析课程提交作业时遇到了这个棘手的问题:本地运行Jupyter Notebook代码能得到正确结果,但Autograder检测时却抛出KeyError: 'STNAME'异常,错误栈明确指向groupby('STNAME')这一行。
完整错误日志:
--------------------------------------------------------------------- KeyError Traceback (most recent call last) /opt/conda/lib/python3.6/site-packages/pandas/indexes/base.py in get_loc(self, key, method, tolerance) 2133 try: -> 2134 return self._engine.get_loc(key) 2135 except KeyError: pandas/index.pyx in pandas.index.IndexEngine.get_loc (pandas/index.c:4433)() pandas/index.pyx in pandas.index.IndexEngine.get_loc (pandas/index.c:4279)() pandas/src/hashtable_class_helper.pxi in pandas.hashtable.PyObjectHashTable.get_item (pandas/hashtable.c:13742)() pandas/src/hashtable_class_helper.pxi in pandas.hashtable.PyObjectHashTable.get_item (pandas/hashtable.c:13696)() KeyError: 'STNAME' During handling of the above exception, another exception occurred: KeyError Traceback (most recent call last) <ipython-input-12-0bb5f5883245> in <module>() ----> 1 answer_six() <ipython-input-9-63797fbac390> in answer_six() 23 24 #group table by stname and census, sorting only the 3 biggest counties population ---> 25 ccensus_groupby_state= (ccensus_df.groupby('STNAME') ['CENSUS2010POP'].nlargest(3) ) 26 #print (ccensus_groupby_state) 27 ft=(ccensus_groupby_state.reset_index()) /opt/conda/lib/python3.6/site-packages/pandas/core/generic.py in groupby(self, by, axis, level, as_index, sort, group_keys, squeeze, **kwargs) 3989 return groupby(self, by=by, axis=axis, level=level, as_index=as_index, 3990 sort=sort, group_keys=group_keys, squeeze=squeeze, -> 3991 **kwargs) 3992 3993 def asfreq(self, freq, method=None, how=None, normalize=False): /opt/conda/lib/python3.6/site-packages/pandas/core/groupby.py in groupby(obj, by, **kwds) 1509 raise TypeError('invalid type: %s' % type(obj)) 1510 -> 1511 return klass(obj, by, **kwds) 1512 1513 /opt/conda/lib/python3.6/site-packages/pandas/core/groupby.py in __init__(self, obj, keys, axis, level, grouper, exclusions, selection, as_index, sort, group_keys, squeeze, **kwargs) 368 level=level, 369 sort=sort, --> 370 mutated=self.mutated) 371 372 self.obj = obj /opt/conda/lib/python3.6/site-packages/pandas/core/groupby.py in _get_grouper(obj, key, axis, level, sort, mutated) 2460 2461 elif is_in_axis(gpr): # df.groupby('name') -> 2462 in_axis, name, gpr = True, gpr, obj[gpr] 2463 exclusions.append(name) 2464 elif isinstance(gpr, Grouper) and gpr.key is not None: /opt/conda/lib/python3.6/site-packages/pandas/core/frame.py in __getitem__(self, key) 2057 return self._getitem_multilevel(key) 2058 else: -> 2059 return self._getitem_column(key) 2060 2061 def _getitem_column(self, key): /opt/conda/lib/python3.6/site-packages/pandas/core/frame.py in _getitem_column(self, key) 2064 # get column 2065 if self.columns.is_unique: -> 2066 return self._get_item_cache(key) 2067 2068 # duplicate columns & possible reduce dimensionality /opt/conda/lib/python3.6/site-packages/pandas/core/generic.py in _get_item_cache(self, item) 1384 res = cache.get(item) 1385 if res is None: -> 1386 values = self._data.get(item) 1387 res = self._box_item_values(item, values) 1388 cache[item] = res /opt/conda/lib/python3.6/site-packages/pandas/core/internals.py in get(self, item, fastpath) 3541 3542 if not isnull(item): -> 3543 loc = self.items.get_loc(item) 3544 else: 3545 indexer = np.arange(len(self.items)) [isnull(self.items)] /opt/conda/lib/python3.6/site-packages/pandas/indexes/base.py in get_loc(self, key, method, tolerance) 2134 return self._engine.get_loc(key) 2135 except KeyError: -> 2136 return self._engine.get_loc(self._maybe_cast_indexer(key)) 2137 2138 indexer = self.get_indexer([key], method=method, tolerance=tolerance) pandas/index.pyx in pandas.index.IndexEngine.get_loc (pandas/index.c:4433)() pandas/index.pyx in pandas.index.IndexEngine.get_loc (pandas/index.c:4279)() pandas/src/hashtable_class_helper.pxi in pandas.hashtable.PyObjectHashTable.get_item (pandas/hashtable.c:13742)() pandas/src/hashtable_class_helper.pxi in pandas.hashtable.PyObjectHashTable.get_item (pandas/hashtable.c:13696)() KeyError: 'STNAME'
针对性解决方案:
排查数据集差异:Autograder使用的数据集可能和你本地版本不一致,比如列名存在大小写偏差(比如本地是
stname但Autograder里是STNAME),或者列名前后带有隐藏空格。建议在代码中添加一行print(ccensus_df.columns)输出所有列名,确认STNAME的存在及格式。如果有空格,可用ccensus_df.columns = ccensus_df.columns.str.strip()统一清理列名。强制按顺序执行Notebook:Jupyter允许乱序执行单元格,但Autograder会严格从上到下执行所有内容。如果你的数据加载/预处理单元格没有先于
answer_six()执行,可能导致ccensus_df未正确初始化,缺失STNAME列。建议重启内核后,按顺序执行所有单元格,确保数据框在groupby前已完整加载。检查预处理代码的误操作:确认
answer_six()之前的代码中,有没有不小心删除或重命名STNAME列的操作,比如drop('STNAME', axis=1)或rename(columns={'STNAME': 'OtherName'})——这类操作可能在本地测试时被注释,但提交代码时遗漏了恢复。添加列存在性校验:在
groupby前加入断言代码,提前暴露问题:assert 'STNAME' in ccensus_df.columns, "Error: STNAME column is missing from the DataFrame!"
内容的提问来源于stack exchange,提问作者l0rd-r4yden

