Python Pandas代码在Ubuntu虚拟机出现KeyError问题求助:无法找到'questions'列
问题分析与解决
嘿,这个问题其实是个很容易忽略的拼写小失误!咱们来一步步拆解:
核心错误原因
你在代码里明确把DataFrame的列名设置成了:
df.columns=["question","answers"]
这里的列名是单数的question,但在get_cleaned_sentences函数里,你却一直在引用复数的questions:
sents=df[["questions"]]; # 还有下面这行 cleaned=clean_sentence(row["questions"], stopwords);
这就导致Ubuntu环境里的代码找不到对应的列,直接抛出KeyError。
为什么其他环境能正常运行?
大概率是这两种情况之一:
- 你在Colab和Windows环境中使用的
Book11.csv本身就带有questions列,覆盖了你设置的列名; - 你在其他环境运行的是修正过拼写的代码版本,只是自己没注意到差异。
修复方案
只需要把函数里所有的questions改成question就行,修正后的函数代码如下:
def get_cleaned_sentences (df, stopwords=False): sents=df[["question"]]; # 这里改单数 cleaned_sentences=[] for index,row in df.iterrows(): cleaned=clean_sentence(row["question"], stopwords); # 这里也改单数 cleaned_sentences.append(cleaned); return cleaned_sentences;
额外排查建议
以后遇到这类跨环境的列名问题,建议在加载CSV后立刻打印列名确认:
df=pd.read_csv("Book11.csv", encoding= 'cp1252'); df.columns=["question","answers"] print("当前DataFrame列名:", df.columns) # 加这行快速验证
这样能第一时间确认列名是否符合预期,避免这类拼写乌龙。
内容的提问来源于stack exchange,提问作者Alexios
相关产品推荐
相关产品推荐

