You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python机器学习预测Y_Test时出现KeyError: action_taken求助

Fixing the KeyError: action_taken in Your ML Prediction Workflow

Hey there! Let's break down why you're hitting that frustrating KeyError and get your decision tree prediction workflow back on track.

The Root Cause of Your Error

The issue boils down to a simple misstep in how you filtered your training data:

  • You defined feature_variable as a list of columns to use as model features, but this list doesn't include your target column action_taken.
  • Then you ran Training_logs = Training_logs[feature_variable] — this removes every column from your training DataFrame except the ones in feature_variable, including action_taken. Later when you try to access Training_logs['action_taken'], Python can't find it because you already deleted it!

Corrected Code Walkthrough

Here's the fixed version of your code with key changes explained:

import pandas as pd #provide dataframe format
import numpy as np #support all high-level mathematical functions
from sklearn import tree #provides various ML feature such as various classification, regression and clustering algorithms
from sklearn.metrics import confusion_matrix
from sklearn.metrics import classification_report
from sklearn.tree import DecisionTreeClassifier
from sklearn.model_selection import train_test_split
from matplotlib import pyplot as plt #plotting library for python
import seaborn as sns #provide high level informative statistical and attractive graphics
from sklearn import metrics # for accuracy score
from sklearn.tree import export_graphviz
from sklearn.externals.six import StringIO
from IPython.display import Image
import pydotplus
from scipy import misc
%matplotlib inline

# Load and clean raw data
Training_logs=pd.read_csv('C:/Users/path/PA_training_data_reformed_2.CSV')
Training_logs=Training_logs.fillna('')
Testing_logs=pd.read_csv('C:/Users/path/PA_testing_data_reformed_2.csv')
Testing_logs= Testing_logs.fillna('')

# Define categorical value lists (unchanged)
corrected_data=['deny','drop','reset-client','reset-server','reset-both','block-url','block-ip','random-drop','sinkhole','syncookie-sent','block-continue','continue','block-overide','override-lockout','override']
severity= ['medium','high','critical']
action=['alert','allow','corrected_data']
event_category=['THREAT','CORRELATION']
sub_category=['spyware','url','virus','vulnerability','wildfire','wildfire-virus']
traffic_direction=['NS','EW','SN']

# Define feature columns (keep target column OUT of this list!)
feature_variable=['severity','event_category', 'action', 'sub_category','traffic_direction']

# --------------------------
# Key Fix 1: Separate features and target BEFORE filtering
# --------------------------
# Extract features from training data (keep target column intact in original df)
X_train_raw = Training_logs[feature_variable]
# Extract target variable directly from original training data
y_train_raw = Training_logs['action_taken']

# Process features with DictVectorizer
from sklearn.feature_extraction import DictVectorizer
from sklearn.preprocessing import LabelEncoder

X_dict = X_train_raw.to_dict(orient='records')
vect = DictVectorizer(sparse=False)
X_vector = vect.fit_transform(X_dict)

# Encode target variable
le = LabelEncoder()
y_train = le.fit_transform(y_train_raw[:-1])
X_train = X_vector[:-1]

# --------------------------
# Key Fix 2: Correct test data processing
# Use the fitted vectorizer from training to transform test data
# --------------------------
X_test_raw = Testing_logs[feature_variable]
X_test_dict = X_test_raw.to_dict(orient='records')
X_test = vect.transform(X_test_dict)

# Now train your model and make predictions
clf = DecisionTreeClassifier()
clf.fit(X_train, y_train)
y_pred = clf.predict(X_test)

# Optional: If your test data has action_taken for evaluation, uncomment below
# y_test = le.transform(Testing_logs['action_taken'])
# print(classification_report(y_test, y_pred))
# print(confusion_matrix(y_test, y_pred))

Quick Best Practices to Avoid This in the Future

  • Always separate features and target explicitly: Never filter your entire dataset down to just features before extracting your target column.
  • Reuse fitted transformers on test data: Always use the vectorizer/encoder you fitted on training data to process test data — don't refit or reuse training data vectors for testing.
  • Double-check column existence: Before accessing any column, you can run print(Training_logs.columns) to confirm it's still present in your DataFrame.

内容的提问来源于stack exchange,提问作者Murtuza Husain

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:53:51