使用Surprise库构建推荐系统时遭遇ZeroDivisionError问题的解决咨询(无法丢弃0值场景)
Let's break down why you're hitting this ZeroDivisionError and how to fix it, even with binary ratings (0/1) that you can't discard.
Root Cause
The error happens because Surprise's NMF implementation uses stochastic gradient descent (SGD) updates that can hit division-by-zero when dealing with extreme rating distributions—like your dataset has duplicate entries (user a has 3 identical q1/False entries) and a small, sparse dataset. When updating the user/item matrices during training, the math can lead to a zero denominator if there's not enough varied data to stabilize the calculations.
Step-by-Step Fixes
1. Remove Duplicate Entries First
Your dataset has redundant rows (e.g., three identical ['a', 'q1', False] entries). These duplicates skew the rating distribution and make the model's updates unstable. Let's clean that up:
# Drop duplicate user-question pairs to balance the dataset df_testing = df_testing.drop_duplicates(subset=['user', 'question'], keep='first')
2. Adjust NMF Parameters to Avoid Extreme Values
Surprise's NMF has regularization parameters that can prevent the SGD updates from hitting zero denominators. Add reg_pu (user matrix regularization) and reg_qi (item matrix regularization) to stabilize training, plus increase epochs for better convergence:
algo = NMF(reg_pu=0.05, reg_qi=0.05, n_epochs=50)
3. Try Alternative Algorithms (If NMF Still Fails)
If you still run into issues, some Surprise algorithms handle binary/sparse data better out of the box:
- SVD: More robust to small datasets and binary ratings
- KNNWithMeans: Accounts for average ratings to smooth out extreme values
- CoClustering: Works well for sparse, categorical-like rating data
For example, switching to SVD would look like this:
algo = SVD(n_epochs=50, reg_all=0.05)
Full Working Code
Here's the revised code incorporating all fixes, plus a small typo correction from your original code:
from surprise import NMF, SVD, Reader, Dataset import numpy as np import pandas as pd # Original dataset a = [['a', 'q1', False], ['a', 'q1', False], ['a', 'q1', False], ['b', 'q1', True], ['b', 'q2', True], ['c', 'q1', True], ['c', 'q2', True], ['c', 'q3', False], ['c', 'q4', False], ['d', 'q1', False], ['e', 'q1', True], ['e', 'q2', True], ['e', 'q3', True], ['e', 'q4', True], ['e', 'q5', True], ['f', 'q3', True], ['f', 'q4', True], ['f', 'q5', True], ['f', 'q6', True], ['g', 'q1', True], ['g', 'q2', True]] df_testing = pd.DataFrame(a, columns=['user', 'question', 'Truth']) # Step 1: Remove duplicate user-question pairs df_testing = df_testing.drop_duplicates(subset=['user', 'question'], keep='first') # Convert boolean ratings to 0/1 integers for Surprise df_testing['Truth'] = df_testing['Truth'].astype(int) # Load data into Surprise-compatible format reader = Reader(rating_scale=(0, 1)) data = Dataset.load_from_df(df_testing[['user', 'question', 'Truth']], reader) # Identify questions user 'f' hasn't interacted with yet unique_qs = df_testing['question'].unique() user_f_qs = df_testing.loc[df_testing['user']=='f', 'question'] qs_to_predict = np.setdiff1d(unique_qs, user_f_qs) # Step 2: Use regularized NMF (swap with SVD if needed) algo = NMF(reg_pu=0.05, reg_qi=0.05, n_epochs=50) # algo = SVD(n_epochs=50, reg_all=0.05) # Alternative robust option # Train the model and generate predictions algo.fit(data.build_full_trainset()) my_recs = [] for q_id in qs_to_predict: # Note: Surprise uses 'iid' for item IDs, not '_id' (fixed typo here) pred_score = algo.predict(uId='f', iid=q_id).est my_recs.append((q_id, pred_score)) # Sort and display recommendations recommendations = pd.DataFrame(my_recs, columns=['question', 'prediction']).sort_values('prediction', ascending=False) print(recommendations)
Key Notes
- I fixed a small typo in your original code: Surprise's
predictmethod usesiidfor item IDs, not_id. - Regularization parameters add small penalties to prevent the model from overfitting and hitting extreme values that cause division-by-zero.
- Removing duplicates is critical here—your original dataset had 3 identical entries for user
aandq1, which made the rating distribution unbalanced and destabilized training.
内容的提问来源于stack exchange,提问作者futuredataengineer

