使用Scikit-Learn KNeighborsClassifier解决特征维度不匹配问题
Got it, let's tackle this problem step by step. The ValueError you're seeing is straightforward: your KNeighborsClassifier was trained on 18 statistical features (goals, fouls, etc.), but your test data only provides two team names—so the feature dimensions don't match. KNN relies on calculating distance between data points, which requires identical input dimensions. Here are practical, actionable fixes that preserve both your historical training data and your limited test features:
1. Map Team Names to Their Historical Statistical Profiles
The most direct fix is to convert the two team names in your test set into the same 18-dimensional feature space your model was trained on. You'll do this by leveraging the historical stats from your training data to build a "profile" for each team.
Step-by-Step Implementation:
- First, build a team-to-stats dictionary using your training data. Calculate the average of each of the 18 features for every team (combining their home and away game stats to get a robust profile):
import pandas as pd # Assume your training data is in a DataFrame called train_df # Define your 18 statistical feature columns (adjust to match your data) stats_columns = ["goals", "fouls", "shots_on_target", ...] # total 18 columns # Combine home and away team data to calculate overall averages per team home_team_data = train_df[["home_team"] + stats_columns].rename(columns={"home_team": "team"}) away_team_data = train_df[["away_team"] + stats_columns].rename(columns={"away_team": "team"}) all_team_stats = pd.concat([home_team_data, away_team_data]) # Calculate average stats for each team team_stat_profiles = all_team_stats.groupby("team")[stats_columns].mean().to_dict("index") - Transform your test set by replacing each team name with their precomputed 18-dimensional stats. If your training data's features include separate home/away stats (e.g.,
home_goals,away_goals), make sure to align the test features to match this structure:# Assume test data has columns "home_team_test" and "away_team_test" test_df = pd.DataFrame({"home_team_test": ["TeamA"], "away_team_test": ["TeamB"]}) # Map home team to its stats, then away team to its stats test_home_stats = test_df["home_team_test"].map(lambda x: team_stat_profiles[x]) test_away_stats = test_df["away_team_test"].map(lambda x: team_stat_profiles[x]) # Convert to DataFrame and concatenate to get 18 features (match training order!) test_home_df = pd.DataFrame(test_home_stats.tolist(), columns=[f"home_{col}" for col in stats_columns]) test_away_df = pd.DataFrame(test_away_stats.tolist(), columns=[f"away_{col}" for col in stats_columns]) final_test_features = pd.concat([test_home_df, test_away_df], axis=1) - Now
final_test_featureshas the same 18 dimensions as your training data, so you can pass it to your trainedKNeighborsClassifierwithout errors.
Edge Case Handling:
If your test set includes a team that wasn't in the training data, use the league-wide average stats for all 18 features as a fallback.
2. Train a Team Embedding Model for Dynamic Feature Mapping
If teams' performance varies over time (e.g., different seasons), a static average profile might not be ideal. Instead, train a simple embedding model to map team names to a 18-dimensional vector that captures their historical performance patterns.
How to Do It:
- Use your training data to train a dimensionality reduction model (like PCA) on the 18 statistical features. Then, for each team, use their historical stats to generate an embedding:
from sklearn.decomposition import PCA # Fit PCA to the training stats to get a 18-dimensional embedding pca = PCA(n_components=18) pca.fit(train_df[stats_columns]) # For each team, generate their embedding by transforming their average stats team_embeddings = {} for team in team_stat_profiles: team_avg_stats = pd.Series(team_stat_profiles[team]).values.reshape(1, -1) team_embeddings[team] = pca.transform(team_avg_stats)[0] - Use these embeddings to convert your test set's team names into 18-dimensional features, just like in Method 1. This approach adapts to patterns in the training data rather than just using raw averages.
3. Reframe the Model to Use Team Pair Similarity
If you want to lean into the team name input directly, reframe your KNN model to work with team pairs instead of raw stats. Here's how:
- Preprocess your training data to use (home_team, away_team) as a key, and store the corresponding 18 stats and match result.
- For a test pair (TeamX, TeamY), calculate similarity scores between TeamX and all home teams in the training set, and TeamY and all away teams.
- Use these similarity scores to weight the K nearest training samples, then predict based on those weighted results.
This method keeps your historical stats intact while using team names as the primary test input.
内容的提问来源于stack exchange,提问作者duldi

