Python机器学习预测Y_Test时出现KeyError: action_taken求助
Fixing the
KeyError: action_taken in Your ML Prediction Workflow Hey there! Let's break down why you're hitting that frustrating KeyError and get your decision tree prediction workflow back on track.
The Root Cause of Your Error
The issue boils down to a simple misstep in how you filtered your training data:
- You defined
feature_variableas a list of columns to use as model features, but this list doesn't include your target columnaction_taken. - Then you ran
Training_logs = Training_logs[feature_variable]— this removes every column from your training DataFrame except the ones infeature_variable, includingaction_taken. Later when you try to accessTraining_logs['action_taken'], Python can't find it because you already deleted it!
Corrected Code Walkthrough
Here's the fixed version of your code with key changes explained:
import pandas as pd #provide dataframe format import numpy as np #support all high-level mathematical functions from sklearn import tree #provides various ML feature such as various classification, regression and clustering algorithms from sklearn.metrics import confusion_matrix from sklearn.metrics import classification_report from sklearn.tree import DecisionTreeClassifier from sklearn.model_selection import train_test_split from matplotlib import pyplot as plt #plotting library for python import seaborn as sns #provide high level informative statistical and attractive graphics from sklearn import metrics # for accuracy score from sklearn.tree import export_graphviz from sklearn.externals.six import StringIO from IPython.display import Image import pydotplus from scipy import misc %matplotlib inline # Load and clean raw data Training_logs=pd.read_csv('C:/Users/path/PA_training_data_reformed_2.CSV') Training_logs=Training_logs.fillna('') Testing_logs=pd.read_csv('C:/Users/path/PA_testing_data_reformed_2.csv') Testing_logs= Testing_logs.fillna('') # Define categorical value lists (unchanged) corrected_data=['deny','drop','reset-client','reset-server','reset-both','block-url','block-ip','random-drop','sinkhole','syncookie-sent','block-continue','continue','block-overide','override-lockout','override'] severity= ['medium','high','critical'] action=['alert','allow','corrected_data'] event_category=['THREAT','CORRELATION'] sub_category=['spyware','url','virus','vulnerability','wildfire','wildfire-virus'] traffic_direction=['NS','EW','SN'] # Define feature columns (keep target column OUT of this list!) feature_variable=['severity','event_category', 'action', 'sub_category','traffic_direction'] # -------------------------- # Key Fix 1: Separate features and target BEFORE filtering # -------------------------- # Extract features from training data (keep target column intact in original df) X_train_raw = Training_logs[feature_variable] # Extract target variable directly from original training data y_train_raw = Training_logs['action_taken'] # Process features with DictVectorizer from sklearn.feature_extraction import DictVectorizer from sklearn.preprocessing import LabelEncoder X_dict = X_train_raw.to_dict(orient='records') vect = DictVectorizer(sparse=False) X_vector = vect.fit_transform(X_dict) # Encode target variable le = LabelEncoder() y_train = le.fit_transform(y_train_raw[:-1]) X_train = X_vector[:-1] # -------------------------- # Key Fix 2: Correct test data processing # Use the fitted vectorizer from training to transform test data # -------------------------- X_test_raw = Testing_logs[feature_variable] X_test_dict = X_test_raw.to_dict(orient='records') X_test = vect.transform(X_test_dict) # Now train your model and make predictions clf = DecisionTreeClassifier() clf.fit(X_train, y_train) y_pred = clf.predict(X_test) # Optional: If your test data has action_taken for evaluation, uncomment below # y_test = le.transform(Testing_logs['action_taken']) # print(classification_report(y_test, y_pred)) # print(confusion_matrix(y_test, y_pred))
Quick Best Practices to Avoid This in the Future
- Always separate features and target explicitly: Never filter your entire dataset down to just features before extracting your target column.
- Reuse fitted transformers on test data: Always use the vectorizer/encoder you fitted on training data to process test data — don't refit or reuse training data vectors for testing.
- Double-check column existence: Before accessing any column, you can run
print(Training_logs.columns)to confirm it's still present in your DataFrame.
内容的提问来源于stack exchange,提问作者Murtuza Husain
相关产品推荐
相关产品推荐

