代码在Mac终端与VS Code正常运行,Coursera Jupyter Notebook报错
Hey there! Let's break down why you're hitting this KeyError in Coursera's Jupyter Notebook, even though your code works elsewhere.
The problem comes down to how you're handling the STNAME column after setting it as the DataFrame index:
- On line 6, you run
census = census.set_index(['STNAME'])— this movesSTNAMEfrom the regular columns into the index of the DataFrame. - But then on line 7, you try to sort using
['STNAME', 'CENSUS2010POP']insort_values— pandas can't findSTNAMEin the columns anymore, so it throws theKeyError. - The same issue would hit line 9 when you try to group by
['STNAME']for the sum operation.
Why does this work locally? Likely a pandas version difference. Older pandas versions were more lenient about accepting index names in sort_values, but the version in Coursera's environment (paired with Python 3.6, a relatively strict pandas release) enforces the clear distinction between columns and indexes.
Here are two straightforward fixes:
Option 1: Reorder Operations (Sort Before Setting Index)
This is the simplest fix — just sort your data while STNAME is still a regular column, then set it as the index:
import pandas as pd census_df = pd.read_csv('census.csv') def answer_six(): census = census_df[census_df['SUMLEV']==50] colstokeep = ['STNAME', 'CTYNAME', 'CENSUS2010POP'] census = census[colstokeep] # Sort first while STNAME is still a column census = census.sort_values(['STNAME', 'CENSUS2010POP'], ascending=(True, False)) # Now set STNAME as index census = census.set_index(['STNAME']) # Group by index level to get top 3 counties per state census = census.groupby(level=0).head(3) # Sum by index level final = census.groupby(level=0).sum() final = final.sort_values(['CENSUS2010POP'], ascending=False) final_indexes = final.index.values.tolist() return final_indexes[:3] answer_six()
Option 2: Explicitly Use Index in Sort/Group Operations
If you prefer to keep the index set first, adjust the sort_values and groupby calls to reference the index instead of columns:
import pandas as pd census_df = pd.read_csv('census.csv') def answer_six(): census = census_df[census_df['SUMLEV']==50] colstokeep = ['STNAME', 'CTYNAME', 'CENSUS2010POP'] census = census[colstokeep] census = census.set_index(['STNAME']) # Sort by index (STNAME) first, then by population descending census = census.sort_index(level=0, ascending=True).sort_values('CENSUS2010POP', ascending=False) # Group by index level census = census.groupby(level=0).head(3) # Sum by index level final = census.groupby(level=0).sum() final = final.sort_values(['CENSUS2010POP'], ascending=False) final_indexes = final.index.values.tolist() return final_indexes[:3] answer_six()
You can verify the pandas version difference by running print(pd.__version__) in both your local environment and Coursera's notebook — this will confirm if version compatibility is the reason for the differing behavior.
内容的提问来源于stack exchange,提问作者dataadina

