Python Pandas读取文件选列绘图时出现KeyError的问题求解及原因分析
修复选择列后的KeyError问题
错误原因分析
你遇到的KeyError: "None of [Index([('PassengerId', 'Survived')], dtype='object')] are in the [columns]"错误,根源在build_graph()函数里的这行代码:
df = file[[column_list]].plot()
这里的column_list已经是一个包含列名的列表(比如['PassengerId', 'Survived']),但你又给它套了一层方括号,变成了嵌套列表[['PassengerId', 'Survived']]。Pandas会把这个嵌套列表当作一个整体的列索引去查找,而你的数据集里根本不存在名为('PassengerId', 'Survived')的列,所以直接触发了KeyError。
另外还有两个小问题需要注意:
- 你的需求是读取Excel文件,但当前代码用的是
pd.read_csv(),这会导致读取.xlsx文件失败,需要改成pd.read_excel()。 choose_file()里的open(filepath, 'r+')完全没必要,pandas读取文件时会自行处理文件流,手动打开反而可能引发权限或文件占用问题。
修复后的完整代码
import pandas as pd import os import plotly.express as px # 让用户输入文件路径并验证文件存在 def choose_file(): global filepath filepath = input('Enter filepath: ') assert os.path.exists(filepath), "I did not find the file at, " + str(filepath) print("Hooray we found your file!") # 读取Excel文件为Pandas DataFrame并打印列名 def open_file(): global file # 替换read_csv为read_excel,适配Excel文件 file = pd.read_excel(filepath) print("Available columns:") for col in file.columns: print(f"- {col}") # 让用户选择列并存入列表,筛选DataFrame的列 def choose_columns(): global column_list global file column_list = [] needs_items = True while needs_items: user_input = input('Enter a column name to select: ') # 增加列名验证,避免用户输入不存在的列 if user_input in file.columns: column_list.append(user_input) print("Selected columns so far:") for col in column_list: print(f"- {col}") else: print(f"Error: '{user_input}' is not a valid column name. Please try again.") answer = input("Add another column? (y/n) ").lower() if answer == "n": needs_items = False print('Final selected columns: ', column_list) file = file[column_list] print("\nFiltered data preview:") print(file.head()) # 生成图表(这里用Plotly生成交互式图表,更符合需求) def build_graph(): global file # 修复嵌套列表问题,直接用column_list作为索引 if len(column_list) == 2: # 两列数据用散点图展示 fig = px.scatter(file, x=column_list[0], y=column_list[1]) fig.show() elif len(column_list) > 2: # 多列数据用折线图展示 fig = px.line(file) fig.show() else: print("Need at least 2 columns to generate a meaningful chart.") choose_file() open_file() choose_columns() build_graph()
关键修复点说明
- 修复KeyError:把
file[[column_list]]改成file[column_list],直接用已有的列名列表去筛选DataFrame,不需要额外嵌套。 - 适配Excel读取:将
pd.read_csv()替换为pd.read_excel(),满足读取Excel文件的需求。 - 增加列名验证:在
choose_columns()里添加了输入列名的合法性检查,避免用户输入不存在的列导致后续报错。 - 优化图表生成:改用你导入的Plotly库生成交互式图表,比Pandas默认的matplotlib图表更美观且交互性更强,同时根据选择的列数自动选择合适的图表类型。
内容的提问来源于stack exchange,提问作者Moon
相关产品推荐
相关产品推荐

