使用sklearn的BernoulliNB处理丹佛犯罪数据集无有效预测结果求助
Hey there, let's dig into why your BernoulliNB model isn't performing well with the Denver crime dataset. I've worked with similar structured datasets before, so here are some common pitfalls and actionable fixes to try out:
BernoulliNB is built for binary/boolean features (0s and 1s)—but your raw dataset has a mix of unique IDs, categorical strings, and unprocessed dates. Feeding this raw data directly is almost certainly causing poor performance. Let's break down each field type:
- Unique IDs (INCIDENT_ID, OFFENSE_ID, etc.): These are just identifiers with zero predictive value—drop them entirely from your feature set.
- Categorical fields (OFFENSE_TYPE_ID, OFFENSE_CATEGORY_ID): You need to encode these into binary features. Use
OneHotEncoder(for low-cardinality categories) or target encoding (if you have dozens/hundreds of categories). If your target is one of these categorical fields, ensure it's properly encoded too (e.g.,LabelEncoderfor multi-class tasks). - Date fields (FIRST_OCCURRENCE_DATE, etc.): Raw date strings are useless to the model. Extract meaningful binary features like:
is_night(1 if incident is between 10PM-6AM, 0 otherwise)is_weekend(1 if incident falls on Sat/Sun, 0 otherwise)is_peak_hour(1 if during commuting times, 0 otherwise)
For numerical date-derived features (like time between occurrence and report), you’ll need to binarize them (using thebinarizeparameter in BernoulliNB) or switch to a different Naive Bayes variant.
Crime datasets are almost always imbalanced—some offenses (like theft) are way more common than others (like homicide). BernoulliNB assumes equal class priors by default, which can skew predictions toward majority classes.
- Adjust the
class_priorparameter to match the actual distribution of your target classes. - Try oversampling minority classes (e.g., SMOTE) or undersampling majority classes to balance the dataset.
- Use stratified train-test splits to ensure both sets have the same class distribution as the original data. If you're doing temporal prediction (predicting future crimes), use a time-based split instead of random to avoid data leakage.
Small tweaks to the model's parameters can make a big difference:
alpha(smoothing parameter): If your model is overfitting (great train accuracy, bad test accuracy), increase alpha. If it's underfitting, decrease it. UseGridSearchCVto find the optimal value:from sklearn.model_selection import GridSearchCV from sklearn.naive_bayes import BernoulliNB param_grid = {'alpha': [0.1, 0.5, 1.0, 2.0, 5.0]} grid_search = GridSearchCV(BernoulliNB(), param_grid, cv=5, scoring='accuracy') grid_search.fit(X_train_processed, y_train) print(f"Best alpha value: {grid_search.best_params_['alpha']}")binarize: If you have numerical features (like time differences), set a threshold to convert them to binary. The default is 0, but you might need to adjust based on your data (e.g., binarizing "time since last incident" into "within 24h" vs "longer than 24h").
Your LAST_OCCURRENCE_DATE has over 290k nulls. Instead of just dropping these rows, turn missingness into a feature: create a binary column has_last_occurrence (1 if the date exists, 0 if it's null)—missingness might correlate with certain crime types (e.g., one-time incidents vs ongoing crimes).
BernoulliNB shines for text classification or strict binary feature tasks. If your target is multi-class (like predicting crime category) or you have a lot of numerical features, other Naive Bayes variants might perform better:
- MultinomialNB: Works well with count-based features (e.g., frequency of crimes in certain neighborhoods).
- GaussianNB: Better suited for continuous numerical features (e.g., geographic coordinates, time differences).
Here's a snippet to kickstart your workflow:
import pandas as pd from sklearn.preprocessing import OneHotEncoder, FunctionTransformer from sklearn.compose import ColumnTransformer from sklearn.naive_bayes import BernoulliNB from sklearn.model_selection import train_test_split # Load dataset df = pd.read_csv('denver_crime_data.csv') # Drop useless ID columns df = df.drop(['INCIDENT_ID', 'OFFENSE_ID', 'OFFENSE_CODE', 'OFFENSE_CODE_EXTENSION'], axis=1) # Process date features into binary flags def process_dates(X): X['FIRST_OCCURRENCE_DATE'] = pd.to_datetime(X['FIRST_OCCURRENCE_DATE']) X['is_night'] = (X['FIRST_OCCURRENCE_DATE'].dt.hour >= 22) | (X['FIRST_OCCURRENCE_DATE'].dt.hour <= 6) X['is_weekend'] = X['FIRST_OCCURRENCE_DATE'].dt.dayofweek >= 5 X['has_last_occurrence'] = X['LAST_OCCURRENCE_DATE'].notna().astype(int) return X[['is_night', 'is_weekend', 'has_last_occurrence']] # Build preprocessor preprocessor = ColumnTransformer( transformers=[ ('date_features', FunctionTransformer(process_dates), ['FIRST_OCCURRENCE_DATE', 'LAST_OCCURRENCE_DATE']), ('categorical', OneHotEncoder(sparse_output=False, drop='first'), ['OFFENSE_TYPE_ID']) ]) # Define target and features (adjust target based on your task) X = df.drop('OFFENSE_CATEGORY_ID', axis=1) y = df['OFFENSE_CATEGORY_ID'] # Split data with stratification X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, stratify=y, random_state=42) # Preprocess data X_train_processed = preprocessor.fit_transform(X_train) X_test_processed = preprocessor.transform(X_test) # Train tuned BernoulliNB model model = BernoulliNB(alpha=0.5) model.fit(X_train_processed, y_train) # Evaluate print(f"Train Accuracy: {model.score(X_train_processed, y_train):.2f}") print(f"Test Accuracy: {model.score(X_test_processed, y_test):.2f}")
内容的提问来源于stack exchange,提问作者user8101370

