基于Yelp数据集的餐饮类基于内容的推荐系统:数值与文本特征结合下的余弦相似度(Cosine similarity)实现方案问询
Great question! You’re absolutely on the right track—combining text and numerical features by calculating separate similarity scores and then merging them is a standard, effective approach for hybrid content-based recommenders. Let’s break down how to implement this step-by-step, along with practical tips and alternatives:
Step 1: Process Text Features (Categories & Attributes)
This builds on the DataCamp tutorial, but we’ll refine it to handle Yelp’s specific data:
- Clean the text data:
categoriesis already a comma-separated string (e.g., "Restaurants, Pizza")—use it directly.attributesis typically a JSON string (like{'OutdoorSeating': 'True', 'WiFi': 'Free'}). Parse it into a dictionary, then convert it to a flat string (e.g.,"OutdoorSeating:True WiFi:Free") so the vectorizer can capture these as meaningful features.
- Use TF-IDF instead of CountVectorizer: It penalizes overused categories/attributes, leading to more accurate similarity scores.
- Calculate text similarity: Generate a cosine similarity matrix for the text features.
import json from sklearn.feature_extraction.text import TfidfVectorizer from sklearn.metrics.pairwise import cosine_similarity # Parse and clean attributes business_df['attributes_clean'] = business_df['attributes'].apply( lambda x: ' '.join([f"{k}:{v}" for k, v in json.loads(x).items()]) if x else '' ) # Combine categories and attributes into one text feature business_df['text_features'] = business_df['categories'] + ' ' + business_df['attributes_clean'] # TF-IDF vectorization tfidf = TfidfVectorizer(stop_words='english', ngram_range=(1,2)) tfidf_matrix = tfidf.fit_transform(business_df['text_features']) # Text similarity matrix text_similarity = cosine_similarity(tfidf_matrix)
Step 2: Process Numerical Features
Split these into two groups, as they represent distinct types of similarity:
2.1 Geographic Similarity (Latitude & Longitude)
Euclidean distance doesn’t work well for coordinates (the Earth is a sphere). Use the Haversine formula to calculate actual geographic distance, then convert distance to similarity (closer = more similar):
import numpy as np def haversine(lat1, lon1, lat2, lon2): # Convert degrees to radians lat1, lon1, lat2, lon2 = map(np.radians, [lat1, lon1, lat2, lon2]) # Haversine formula for spherical distance dlat = lat2 - lat1 dlon = lon2 - lon1 a = np.sin(dlat/2)**2 + np.cos(lat1) * np.cos(lat2) * np.sin(dlon/2)**2 c = 2 * np.arcsin(np.sqrt(a)) # Radius of Earth in kilometers (use 3956 for miles) km = 6371 * c return km # Vectorized distance calculation (faster than nested loops) coords = business_df[['latitude', 'longitude']].values geo_distance = np.array([haversine(coords[i,0], coords[i,1], coords[:,0], coords[:,1]) for i in range(len(coords))]) # Convert distance to similarity (1/(1+distance) keeps values between 0 and 1) geo_similarity = 1 / (1 + geo_distance)
2.2 Quality Similarity (Stars & Review Count)
First standardize these features (they’re on very different scales), then calculate cosine similarity:
from sklearn.preprocessing import StandardScaler # Extract and scale quality features quality_features = business_df[['stars', 'review_count']] scaler = StandardScaler() scaled_quality = scaler.fit_transform(quality_features) # Quality similarity matrix quality_similarity = cosine_similarity(scaled_quality)
Step 3: Combine Similarity Scores
Assign weights to each component based on your priorities (e.g., if location matters more, increase the geographic weight). Weights should sum to 1 for interpretability:
# Adjust these weights based on your business needs text_weight = 0.4 geo_weight = 0.3 quality_weight = 0.3 # Combined similarity matrix combined_similarity = (text_weight * text_similarity) + (geo_weight * geo_similarity) + (quality_weight * quality_similarity)
Step 4: Add Heat Filtering & Weighted Rating
Follow the DataCamp approach, adapted for Yelp data:
- For a target restaurant, fetch the top 30 most similar entries using the combined similarity matrix.
- Calculate an IMDB-style weighted rating to balance average stars and review count (avoids favoring obscure high-rated spots).
- Sort by weighted rating and return the top 10 recommendations.
def recommend_similar_restaurants(business_name, top_n=10): # Get index of the target restaurant idx = business_df[business_df['name'] == business_name].index[0] # Get sorted similarity scores (exclude the restaurant itself) sim_scores = sorted(list(enumerate(combined_similarity[idx])), key=lambda x: x[1], reverse=True)[1:31] # Get indices of top 30 candidates restaurant_indices = [i[0] for i in sim_scores] top_candidates = business_df.iloc[restaurant_indices].copy() # Calculate weighted rating parameters C = business_df['stars'].mean() # Average stars across all restaurants m = business_df['review_count'].quantile(0.25) # Minimum reviews to qualify (25th percentile) # Filter out low-review candidates top_candidates = top_candidates[top_candidates['review_count'] >= m] # Compute weighted rating top_candidates['weighted_rating'] = (top_candidates['review_count']/(top_candidates['review_count']+m))*top_candidates['stars'] + (m/(top_candidates['review_count']+m))*C # Return top N sorted by weighted rating return top_candidates.sort_values('weighted_rating', ascending=False).head(top_n)
Alternative: Unified Feature Matrix
If you prefer a more streamlined approach:
- Convert text features to TF-IDF vectors.
- Standardize numerical features (stars, review_count, latitude/longitude).
- Concatenate the TF-IDF matrix with scaled numerical features into one big matrix.
- Calculate cosine similarity on this combined matrix.
Note: Text features are high-dimensional, so numerical features might get overshadowed. You’ll need to adjust weights (e.g., multiply numerical features by a scaling factor) or use PCA to balance the feature space. The separate similarity merging method is more flexible and interpretable.
Key Tips
- Add error handling for malformed
attributesentries (some may beNoneor invalid JSON). - Test different weight combinations to optimize recommendations (e.g., increase text weight if users care most about cuisine).
- For large datasets, use vectorized operations or libraries like
geopyto speed up geographic distance calculations.
内容的提问来源于stack exchange,提问作者Vishal Balaji

