You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Yelp数据集的餐饮类基于内容的推荐系统:数值与文本特征结合下的余弦相似度(Cosine similarity)实现方案问询

Great question! You’re absolutely on the right track—combining text and numerical features by calculating separate similarity scores and then merging them is a standard, effective approach for hybrid content-based recommenders. Let’s break down how to implement this step-by-step, along with practical tips and alternatives:

Step 1: Process Text Features (Categories & Attributes)

This builds on the DataCamp tutorial, but we’ll refine it to handle Yelp’s specific data:

  • Clean the text data:
    • categories is already a comma-separated string (e.g., "Restaurants, Pizza")—use it directly.
    • attributes is typically a JSON string (like {'OutdoorSeating': 'True', 'WiFi': 'Free'}). Parse it into a dictionary, then convert it to a flat string (e.g., "OutdoorSeating:True WiFi:Free") so the vectorizer can capture these as meaningful features.
  • Use TF-IDF instead of CountVectorizer: It penalizes overused categories/attributes, leading to more accurate similarity scores.
  • Calculate text similarity: Generate a cosine similarity matrix for the text features.
import json
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.metrics.pairwise import cosine_similarity

# Parse and clean attributes
business_df['attributes_clean'] = business_df['attributes'].apply(
    lambda x: ' '.join([f"{k}:{v}" for k, v in json.loads(x).items()]) if x else ''
)

# Combine categories and attributes into one text feature
business_df['text_features'] = business_df['categories'] + ' ' + business_df['attributes_clean']

# TF-IDF vectorization
tfidf = TfidfVectorizer(stop_words='english', ngram_range=(1,2))
tfidf_matrix = tfidf.fit_transform(business_df['text_features'])

# Text similarity matrix
text_similarity = cosine_similarity(tfidf_matrix)

Step 2: Process Numerical Features

Split these into two groups, as they represent distinct types of similarity:

2.1 Geographic Similarity (Latitude & Longitude)

Euclidean distance doesn’t work well for coordinates (the Earth is a sphere). Use the Haversine formula to calculate actual geographic distance, then convert distance to similarity (closer = more similar):

import numpy as np

def haversine(lat1, lon1, lat2, lon2):
    # Convert degrees to radians
    lat1, lon1, lat2, lon2 = map(np.radians, [lat1, lon1, lat2, lon2])
    # Haversine formula for spherical distance
    dlat = lat2 - lat1
    dlon = lon2 - lon1
    a = np.sin(dlat/2)**2 + np.cos(lat1) * np.cos(lat2) * np.sin(dlon/2)**2
    c = 2 * np.arcsin(np.sqrt(a))
    # Radius of Earth in kilometers (use 3956 for miles)
    km = 6371 * c
    return km

# Vectorized distance calculation (faster than nested loops)
coords = business_df[['latitude', 'longitude']].values
geo_distance = np.array([haversine(coords[i,0], coords[i,1], coords[:,0], coords[:,1]) for i in range(len(coords))])

# Convert distance to similarity (1/(1+distance) keeps values between 0 and 1)
geo_similarity = 1 / (1 + geo_distance)

2.2 Quality Similarity (Stars & Review Count)

First standardize these features (they’re on very different scales), then calculate cosine similarity:

from sklearn.preprocessing import StandardScaler

# Extract and scale quality features
quality_features = business_df[['stars', 'review_count']]
scaler = StandardScaler()
scaled_quality = scaler.fit_transform(quality_features)

# Quality similarity matrix
quality_similarity = cosine_similarity(scaled_quality)

Step 3: Combine Similarity Scores

Assign weights to each component based on your priorities (e.g., if location matters more, increase the geographic weight). Weights should sum to 1 for interpretability:

# Adjust these weights based on your business needs
text_weight = 0.4
geo_weight = 0.3
quality_weight = 0.3

# Combined similarity matrix
combined_similarity = (text_weight * text_similarity) + (geo_weight * geo_similarity) + (quality_weight * quality_similarity)

Step 4: Add Heat Filtering & Weighted Rating

Follow the DataCamp approach, adapted for Yelp data:

  1. For a target restaurant, fetch the top 30 most similar entries using the combined similarity matrix.
  2. Calculate an IMDB-style weighted rating to balance average stars and review count (avoids favoring obscure high-rated spots).
  3. Sort by weighted rating and return the top 10 recommendations.
def recommend_similar_restaurants(business_name, top_n=10):
    # Get index of the target restaurant
    idx = business_df[business_df['name'] == business_name].index[0]
    
    # Get sorted similarity scores (exclude the restaurant itself)
    sim_scores = sorted(list(enumerate(combined_similarity[idx])), key=lambda x: x[1], reverse=True)[1:31]
    
    # Get indices of top 30 candidates
    restaurant_indices = [i[0] for i in sim_scores]
    top_candidates = business_df.iloc[restaurant_indices].copy()
    
    # Calculate weighted rating parameters
    C = business_df['stars'].mean()  # Average stars across all restaurants
    m = business_df['review_count'].quantile(0.25)  # Minimum reviews to qualify (25th percentile)
    
    # Filter out low-review candidates
    top_candidates = top_candidates[top_candidates['review_count'] >= m]
    
    # Compute weighted rating
    top_candidates['weighted_rating'] = (top_candidates['review_count']/(top_candidates['review_count']+m))*top_candidates['stars'] + (m/(top_candidates['review_count']+m))*C
    
    # Return top N sorted by weighted rating
    return top_candidates.sort_values('weighted_rating', ascending=False).head(top_n)

Alternative: Unified Feature Matrix

If you prefer a more streamlined approach:

  1. Convert text features to TF-IDF vectors.
  2. Standardize numerical features (stars, review_count, latitude/longitude).
  3. Concatenate the TF-IDF matrix with scaled numerical features into one big matrix.
  4. Calculate cosine similarity on this combined matrix.

Note: Text features are high-dimensional, so numerical features might get overshadowed. You’ll need to adjust weights (e.g., multiply numerical features by a scaling factor) or use PCA to balance the feature space. The separate similarity merging method is more flexible and interpretable.

Key Tips

  • Add error handling for malformed attributes entries (some may be None or invalid JSON).
  • Test different weight combinations to optimize recommendations (e.g., increase text weight if users care most about cuisine).
  • For large datasets, use vectorized operations or libraries like geopy to speed up geographic distance calculations.

内容的提问来源于stack exchange,提问作者Vishal Balaji

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 08:54:10