You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Surprise库构建推荐系统时遭遇ZeroDivisionError问题的解决咨询(无法丢弃0值场景)

Fixing ZeroDivisionError When Using Surprise's NMF for Binary Recommendation Systems

Let's break down why you're hitting this ZeroDivisionError and how to fix it, even with binary ratings (0/1) that you can't discard.

Root Cause

The error happens because Surprise's NMF implementation uses stochastic gradient descent (SGD) updates that can hit division-by-zero when dealing with extreme rating distributions—like your dataset has duplicate entries (user a has 3 identical q1/False entries) and a small, sparse dataset. When updating the user/item matrices during training, the math can lead to a zero denominator if there's not enough varied data to stabilize the calculations.

Step-by-Step Fixes

1. Remove Duplicate Entries First

Your dataset has redundant rows (e.g., three identical ['a', 'q1', False] entries). These duplicates skew the rating distribution and make the model's updates unstable. Let's clean that up:

# Drop duplicate user-question pairs to balance the dataset
df_testing = df_testing.drop_duplicates(subset=['user', 'question'], keep='first')

2. Adjust NMF Parameters to Avoid Extreme Values

Surprise's NMF has regularization parameters that can prevent the SGD updates from hitting zero denominators. Add reg_pu (user matrix regularization) and reg_qi (item matrix regularization) to stabilize training, plus increase epochs for better convergence:

algo = NMF(reg_pu=0.05, reg_qi=0.05, n_epochs=50)

3. Try Alternative Algorithms (If NMF Still Fails)

If you still run into issues, some Surprise algorithms handle binary/sparse data better out of the box:

  • SVD: More robust to small datasets and binary ratings
  • KNNWithMeans: Accounts for average ratings to smooth out extreme values
  • CoClustering: Works well for sparse, categorical-like rating data

For example, switching to SVD would look like this:

algo = SVD(n_epochs=50, reg_all=0.05)

Full Working Code

Here's the revised code incorporating all fixes, plus a small typo correction from your original code:

from surprise import NMF, SVD, Reader, Dataset
import numpy as np
import pandas as pd

# Original dataset
a = [['a', 'q1', False], ['a', 'q1', False], ['a', 'q1', False], ['b', 'q1', True], ['b', 'q2', True], ['c', 'q1', True], ['c', 'q2', True], ['c', 'q3', False], ['c', 'q4', False], ['d', 'q1', False], ['e', 'q1', True], ['e', 'q2', True], ['e', 'q3', True], ['e', 'q4', True], ['e', 'q5', True], ['f', 'q3', True], ['f', 'q4', True], ['f', 'q5', True], ['f', 'q6', True], ['g', 'q1', True], ['g', 'q2', True]]
df_testing = pd.DataFrame(a, columns=['user', 'question', 'Truth'])

# Step 1: Remove duplicate user-question pairs
df_testing = df_testing.drop_duplicates(subset=['user', 'question'], keep='first')

# Convert boolean ratings to 0/1 integers for Surprise
df_testing['Truth'] = df_testing['Truth'].astype(int)

# Load data into Surprise-compatible format
reader = Reader(rating_scale=(0, 1))
data = Dataset.load_from_df(df_testing[['user', 'question', 'Truth']], reader)

# Identify questions user 'f' hasn't interacted with yet
unique_qs = df_testing['question'].unique()
user_f_qs = df_testing.loc[df_testing['user']=='f', 'question']
qs_to_predict = np.setdiff1d(unique_qs, user_f_qs)

# Step 2: Use regularized NMF (swap with SVD if needed)
algo = NMF(reg_pu=0.05, reg_qi=0.05, n_epochs=50)
# algo = SVD(n_epochs=50, reg_all=0.05)  # Alternative robust option

# Train the model and generate predictions
algo.fit(data.build_full_trainset())
my_recs = []
for q_id in qs_to_predict:
    # Note: Surprise uses 'iid' for item IDs, not '_id' (fixed typo here)
    pred_score = algo.predict(uId='f', iid=q_id).est
    my_recs.append((q_id, pred_score))

# Sort and display recommendations
recommendations = pd.DataFrame(my_recs, columns=['question', 'prediction']).sort_values('prediction', ascending=False)
print(recommendations)

Key Notes

  • I fixed a small typo in your original code: Surprise's predict method uses iid for item IDs, not _id.
  • Regularization parameters add small penalties to prevent the model from overfitting and hitting extreme values that cause division-by-zero.
  • Removing duplicates is critical here—your original dataset had 3 identical entries for user a and q1, which made the rating distribution unbalanced and destabilized training.

内容的提问来源于stack exchange,提问作者futuredataengineer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.01 01:52:48