Python程序读取大型CSV时出现KeyError:无法识别'UNDERLYING_LAST'列
解决Black-Scholes模型代码中的KeyError问题
我正在开发Python程序,通过Black-Merton-Scholes方程对比苹果期权的历史价格与预测价格,使用Kaggle上2021-2023年的期权CSV数据集,但运行代码时始终触发KeyError。
代码如下:
from math import log, sqrt, exp from mibian import BS from sklearn.metrics import mean_squared_error import matplotlib.pyplot as plt historical_data = pd.read_csv('historical_data.csv') underlying_price = historical_data['UNDERLYING_LAST'] strike_price = historical_data['STRIKE'] time_to_expiration = historical_data['DTE'] / 365.0 risk_free_rate = 0.05 volatility = historical_data['C_IV'] / 100.0 def black_scholes(price, strike, time, rate, volatility): call_option = BS([price, strike, rate, time], volatility=volatility) return call_option.callPrice theoretical_prices = [black_scholes(price, strike, time, risk_free_rate, volatility) for price, strike, time in zip(underlying_price, strike_price, time_to_expiration)] market_prices = historical_data['C_LAST'] mse = mean_squared_error(market_prices, theoretical_prices) plt.plot(historical_data['QUOTE_DATE'], market_prices, label='Market Price') plt.plot(historical_data['QUOTE_DATE'], theoretical_prices, label='Theoretical Price') plt.xlabel('Date') plt.ylabel('Option Price') plt.title('Black-Scholes Model Performance - Call Option for Apple') plt.legend() plt.show()
报错信息:
Traceback (most recent call last): File "/Users/shashishmac/PythonProjects/GTBS/venv/lib/python3.10/site-packages/pandas/core/indexes/base.py", line 3805, in get_loc return self._engine.get_loc(casted_key) File "index.pyx", line 167, in pandas._libs.index.IndexEngine.get_loc File "index.pyx", line 196, in pandas._libs.index.IndexEngine.get_loc File "pandas/_libs/hashtable_class_helper.pxi", line 7081, in pandas._libs.hashtable.PyObjectHashTable.get_item File "pandas/_libs/hashtable_class_helper.pxi", line 7089, in pandas._libs.hashtable.PyObjectHashTable.get_item KeyError: 'UNDERLYING_LAST' The above exception was the direct cause of the following exception: Traceback (most recent call last): File "/Users/shashishmac/PythonProjects/GTBS/main.py", line 48, in <module> underlying_price = historical_data['UNDERLYING_LAST'] File "/Users/shashishmac/PythonProjects/GTBS/venv/lib/python3.10/site-packages/pandas/core/frame.py", line 4102, in __getitem__ indexer = self.columns.get_loc(key) File "/Users/shashishmac/PythonProjects/GTBS/venv/lib/python3.10/site-packages/pandas/core/indexes/base.py", line 3812, in get_loc raise KeyError(key) from err KeyError: 'UNDERLYING_LAST'
解决步骤:
检查CSV列名匹配性
先打印数据集的所有列名,确认是否存在代码中用到的字段:import pandas as pd historical_data = pd.read_csv('historical_data.csv') print(historical_data.columns.tolist())可能存在列名大小写不一致、拼写错误或带空格的情况,比如实际列名是
Underlying_Last而非UNDERLYING_LAST,需要将代码中的列名修改为与CSV完全一致的名称。补全缺失的pandas导入
原代码使用了pd.read_csv但未导入pandas库,需在代码开头添加:import pandas as pd排查CSV读取异常
如果CSV文件使用了非逗号分隔符或特殊编码,读取时需指定对应参数:historical_data = pd.read_csv('historical_data.csv', sep=';', encoding='utf-8')可根据实际文件格式调整
sep参数(如制表符\t)。批量校验所有用到的列
代码中涉及的列包括UNDERLYING_LAST、STRIKE、DTE、C_IV、C_LAST、QUOTE_DATE,需逐一确认这些列名在数据集中都存在。
内容的提问来源于stack exchange,提问作者Achilles
相关产品推荐
相关产品推荐

