基于Scikit-learn无监督算法实现Bro网络日志全量实时异常检测
Hey there! I’ve worked through similar Bro (now often called Zeek) log anomaly detection workflows, so let’s walk through how to adjust your setup for historical training and real-time inference with scikit-learn’s IsolationForest and KMeans:
历史日志训练的核心准备
First, you’ll need to standardize feature engineering across all your historical Bro logs—unsupervised models rely heavily on consistent input features. Convert structured log fields (like source IP, destination port, byte count, protocol type) into numerical features the model can process:- For discrete fields (e.g., protocol, service type), use
OneHotEncoderorLabelEncoderto convert categorical values to numbers; - For continuous fields (e.g., byte count, session duration), apply
StandardScalerorMinMaxScalerto normalize ranges—this prevents large-value features from dominating the model’s learning.
Feed all preprocessed historical feature data into IsolationForest or KMeans in one batch, so the model learns the baseline pattern of normal network traffic.
- For discrete fields (e.g., protocol, service type), use
实时检测的模型部署要点
Save your trained model and preprocessing transformers (scalers, encoders) usingjobliborpickle—this lets you reuse the exact same preprocessing logic for new logs:- When a new Bro log entry comes in, extract features using the same rules you used for historical data;
- Apply the saved transformers to standardize/encode the new feature vector;
- Run inference with the model:
- For IsolationForest: The model returns
-1for anomalies and1for normal entries—this is a direct binary flag; - For KMeans: Calculate the distance from the new sample to its nearest cluster center. Set a threshold (e.g., the 95th percentile of distances from historical normal samples) and mark any sample exceeding this threshold as anomalous.
- For IsolationForest: The model returns
无监督模型的关键优化注意事项
- Handle concept drift: Network traffic patterns change over time. Schedule periodic retraining with updated historical logs (including any confirmed anomalies you’ve collected) to keep the model aligned with current normal behavior;
- Tune thresholds: For KMeans, use the distance distribution of your historical normal data to set a meaningful threshold (95th or 99th percentile works well for most cases). For IsolationForest, adjust the
contaminationparameter—if you have rough estimates of historical anomaly rates, set it to that value; otherwise start with 0.01-0.05 and tweak based on real-world results; - Prune unnecessary features: Avoid overloading the model with high-cardinality features like raw source IPs. Instead, derive features like "is internal IP" or IP subnet to reduce dimensionality and improve model efficiency.
内容的提问来源于stack exchange,提问作者Haris

