You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用sklearn的BernoulliNB处理丹佛犯罪数据集无有效预测结果求助

Hey there, let's dig into why your BernoulliNB model isn't performing well with the Denver crime dataset. I've worked with similar structured datasets before, so here are some common pitfalls and actionable fixes to try out:

1. Align your features with BernoulliNB's core assumptions

BernoulliNB is built for binary/boolean features (0s and 1s)—but your raw dataset has a mix of unique IDs, categorical strings, and unprocessed dates. Feeding this raw data directly is almost certainly causing poor performance. Let's break down each field type:

  • Unique IDs (INCIDENT_ID, OFFENSE_ID, etc.): These are just identifiers with zero predictive value—drop them entirely from your feature set.
  • Categorical fields (OFFENSE_TYPE_ID, OFFENSE_CATEGORY_ID): You need to encode these into binary features. Use OneHotEncoder (for low-cardinality categories) or target encoding (if you have dozens/hundreds of categories). If your target is one of these categorical fields, ensure it's properly encoded too (e.g., LabelEncoder for multi-class tasks).
  • Date fields (FIRST_OCCURRENCE_DATE, etc.): Raw date strings are useless to the model. Extract meaningful binary features like:
    • is_night (1 if incident is between 10PM-6AM, 0 otherwise)
    • is_weekend (1 if incident falls on Sat/Sun, 0 otherwise)
    • is_peak_hour (1 if during commuting times, 0 otherwise)
      For numerical date-derived features (like time between occurrence and report), you’ll need to binarize them (using the binarize parameter in BernoulliNB) or switch to a different Naive Bayes variant.
2. Fix class imbalance and train-test splitting

Crime datasets are almost always imbalanced—some offenses (like theft) are way more common than others (like homicide). BernoulliNB assumes equal class priors by default, which can skew predictions toward majority classes.

  • Adjust the class_prior parameter to match the actual distribution of your target classes.
  • Try oversampling minority classes (e.g., SMOTE) or undersampling majority classes to balance the dataset.
  • Use stratified train-test splits to ensure both sets have the same class distribution as the original data. If you're doing temporal prediction (predicting future crimes), use a time-based split instead of random to avoid data leakage.
3. Tune BernoulliNB's hyperparameters

Small tweaks to the model's parameters can make a big difference:

  • alpha (smoothing parameter): If your model is overfitting (great train accuracy, bad test accuracy), increase alpha. If it's underfitting, decrease it. Use GridSearchCV to find the optimal value:
    from sklearn.model_selection import GridSearchCV
    from sklearn.naive_bayes import BernoulliNB
    
    param_grid = {'alpha': [0.1, 0.5, 1.0, 2.0, 5.0]}
    grid_search = GridSearchCV(BernoulliNB(), param_grid, cv=5, scoring='accuracy')
    grid_search.fit(X_train_processed, y_train)
    print(f"Best alpha value: {grid_search.best_params_['alpha']}")
    
  • binarize: If you have numerical features (like time differences), set a threshold to convert them to binary. The default is 0, but you might need to adjust based on your data (e.g., binarizing "time since last incident" into "within 24h" vs "longer than 24h").
4. Handle missing values properly

Your LAST_OCCURRENCE_DATE has over 290k nulls. Instead of just dropping these rows, turn missingness into a feature: create a binary column has_last_occurrence (1 if the date exists, 0 if it's null)—missingness might correlate with certain crime types (e.g., one-time incidents vs ongoing crimes).

5. Consider if BernoulliNB is the right fit

BernoulliNB shines for text classification or strict binary feature tasks. If your target is multi-class (like predicting crime category) or you have a lot of numerical features, other Naive Bayes variants might perform better:

  • MultinomialNB: Works well with count-based features (e.g., frequency of crimes in certain neighborhoods).
  • GaussianNB: Better suited for continuous numerical features (e.g., geographic coordinates, time differences).
Quick Preprocessing & Training Example

Here's a snippet to kickstart your workflow:

import pandas as pd
from sklearn.preprocessing import OneHotEncoder, FunctionTransformer
from sklearn.compose import ColumnTransformer
from sklearn.naive_bayes import BernoulliNB
from sklearn.model_selection import train_test_split

# Load dataset
df = pd.read_csv('denver_crime_data.csv')

# Drop useless ID columns
df = df.drop(['INCIDENT_ID', 'OFFENSE_ID', 'OFFENSE_CODE', 'OFFENSE_CODE_EXTENSION'], axis=1)

# Process date features into binary flags
def process_dates(X):
    X['FIRST_OCCURRENCE_DATE'] = pd.to_datetime(X['FIRST_OCCURRENCE_DATE'])
    X['is_night'] = (X['FIRST_OCCURRENCE_DATE'].dt.hour >= 22) | (X['FIRST_OCCURRENCE_DATE'].dt.hour <= 6)
    X['is_weekend'] = X['FIRST_OCCURRENCE_DATE'].dt.dayofweek >= 5
    X['has_last_occurrence'] = X['LAST_OCCURRENCE_DATE'].notna().astype(int)
    return X[['is_night', 'is_weekend', 'has_last_occurrence']]

# Build preprocessor
preprocessor = ColumnTransformer(
    transformers=[
        ('date_features', FunctionTransformer(process_dates), ['FIRST_OCCURRENCE_DATE', 'LAST_OCCURRENCE_DATE']),
        ('categorical', OneHotEncoder(sparse_output=False, drop='first'), ['OFFENSE_TYPE_ID'])
    ])

# Define target and features (adjust target based on your task)
X = df.drop('OFFENSE_CATEGORY_ID', axis=1)
y = df['OFFENSE_CATEGORY_ID']

# Split data with stratification
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, stratify=y, random_state=42)

# Preprocess data
X_train_processed = preprocessor.fit_transform(X_train)
X_test_processed = preprocessor.transform(X_test)

# Train tuned BernoulliNB model
model = BernoulliNB(alpha=0.5)
model.fit(X_train_processed, y_train)

# Evaluate
print(f"Train Accuracy: {model.score(X_train_processed, y_train):.2f}")
print(f"Test Accuracy: {model.score(X_test_processed, y_test):.2f}")

内容的提问来源于stack exchange,提问作者user8101370

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 08:56:01