基于条件评分创建排名分布:求《Babe》高评分用户偏好Top5电影
Got it, let's tackle this problem step by step. I'll assume you're working with a MovieLens-style dataset (super common for movie rating tasks) using pandas—since that's the most typical setup for this kind of analysis. Here's a complete, working solution that fixes the user filtering issue and gets you the top 5 movies these users love:
Step 1: Setup and Load Your Data
First, make sure you have pandas installed, then load your ratings and movies datasets. I'll use standard column names here—adjust if your data uses different labels.
import pandas as pd # Load datasets (replace file paths with your actual data locations) ratings = pd.read_csv('ratings.csv') movies = pd.read_csv('movies.csv')
Step 2: Filter Users Who Rated "Babe" 4 or 5 Stars
This is probably where your code broke earlier. You need to first map the movie title "Babe" to its unique ID, then pull all users who gave it a 4 or 5 star rating.
# Get the movie ID for "Babe" babe_movie_id = movies[movies['title'] == 'Babe']['movie_id'].iloc[0] # Filter users who rated Babe 4 or 5 stars babe_fans = ratings[(ratings['movie_id'] == babe_movie_id) & (ratings['rating'] >= 4)]['user_id'].unique()
Common Pitfalls Here:
- Forgetting that you need to link the movie title to its ID (most rating datasets use IDs instead of titles directly)
- Using
==instead of>=for ratings, or accidentally including lower scores - Not using
.unique()to avoid duplicate user IDs
Step 3: Get All Ratings From Babe Fans (Excluding Babe)
Now we'll pull every rating from these users, but exclude Babe itself so we don't count it in the rankings.
# Get all ratings from Babe fans, excluding Babe babe_fan_ratings = ratings[(ratings['user_id'].isin(babe_fans)) & (ratings['movie_id'] != babe_movie_id)]
Step 4: Calculate Preference Metrics for Other Movies
To get meaningful rankings, we should consider both the average rating and how many Babe fans rated each movie (to avoid one-off high scores from a single user). Let's set a minimum threshold of 10 ratings to ensure relevance.
# Calculate average rating and number of ratings per movie movie_preferences = babe_fan_ratings.groupby('movie_id').agg( avg_rating=('rating', 'mean'), num_ratings=('rating', 'count') ).reset_index() # Filter out movies with too few ratings (adjust threshold as needed) movie_preferences = movie_preferences[movie_preferences['num_ratings'] >= 10] # Merge with movies dataset to get titles movie_preferences = movie_preferences.merge(movies, on='movie_id')
Step 5: Rank and Get Top 5 Movies
Sort by average rating (descending), then number of ratings (descending) to break ties, and pick the top 5.
# Sort by average rating (highest first) and number of ratings (to break ties) ranked_movies = movie_preferences.sort_values(by=['avg_rating', 'num_ratings'], ascending=False) # Get top 5 movies top_5_movies = ranked_movies[['title', 'avg_rating', 'num_ratings']].head(5) # Print the result print("Top 5 Movies Preferred by Users Who Rated 'Babe' 4 or 5 Stars:") print(top_5_movies)
Example Output:
Top 5 Movies Preferred by Users Who Rated 'Babe' 4 or 5 Stars: title avg_rating num_ratings 123 The Godfather 4.85 127 456 Schindler's List 4.82 112 789 Forrest Gump 4.78 156 1011 The Shawshank Redemption 4.75 143 1213 Braveheart 4.70 98
If your dataset has different column names (like userID instead of user_id), just adjust those in the code. Let me know if you hit any specific errors with your actual data!
内容的提问来源于stack exchange,提问作者CodeChallenging

