多分类特征Feature Hashing批量实现:替代手动逐个处理的高效方案
Great question! When you’re dealing with dozens of categorical features (like 40 in your scenario), manually setting up a separate FeatureHasher for each column is tedious and error-prone. Let’s walk through two scalable, clean approaches to solve this problem—plus fix a small typo in your original code first.
Quick Fix for Your Original Code
You initialized fh1 and fh2 but tried to call fh.fit_transform() (where fh isn’t defined) — that’s a typo that would throw an error. The correct calls should use fh1 for Genre and fh2 for Publisher.
Approach 1: Batch Processing with a Loop
This method is straightforward and easy to customize. We’ll loop through all target categorical columns, apply FeatureHasher to each, and combine the results with the original data.
import pandas as pd from sklearn.feature_extraction import FeatureHasher # Sample data (scaled to mimic your 40-feature use case) d = { 'Genre': ['Platform', 'Racing','Sports','Roleplaying','Puzzle','Platform'], 'Publisher': ['Nintendo', 'Noir','Laura','John','John','Noir'], 'Developer': ['DevA', 'DevB', 'DevC', 'DevA', 'DevB', 'DevC'], 'Region': ['NA', 'EU', 'JP', 'NA', 'EU', 'JP'] # Add 36 more categorical columns here } df = pd.DataFrame(data=d) # Configuration: Define which columns to hash and hash size per column target_columns = df.columns.tolist() # Use a subset like ['Genre', 'Publisher', ...] if needed hash_size_per_col = 6 # Start with the original categorical columns processed_dfs = [df[target_columns]] # Loop through each column to generate hashed features for col in target_columns: # Initialize hasher for the column hasher = FeatureHasher(n_features=hash_size_per_col, input_type='string') # Transform and convert to array/DataFrame hashed_array = hasher.fit_transform(df[col]).toarray() # Name columns clearly to avoid duplicate "0/1/2..." labels hashed_df = pd.DataFrame(hashed_array, columns=[f"{col}_hash_{i}" for i in range(hash_size_per_col)]) processed_dfs.append(hashed_df) # Combine all into one final DataFrame final_df = pd.concat(processed_dfs, axis=1) print(final_df.head())
Key Benefits:
- Simple to understand and modify
- Gives you full control over column naming (critical for debugging 40+ features)
- No extra sklearn dependencies beyond
FeatureHasher
Approach 2: Use ColumnTransformer (Recommended for ML Pipelines)
If you plan to integrate this into a machine learning workflow, sklearn.compose.ColumnTransformer is the more elegant solution. It lets you define all hashers in one place and seamlessly combines results.
import pandas as pd from sklearn.feature_extraction import FeatureHasher from sklearn.compose import ColumnTransformer # Same sample data as above d = { 'Genre': ['Platform', 'Racing','Sports','Roleplaying','Puzzle','Platform'], 'Publisher': ['Nintendo', 'Noir','Laura','John','John','Noir'], 'Developer': ['DevA', 'DevB', 'DevC', 'DevA', 'DevB', 'DevC'], 'Region': ['NA', 'EU', 'JP', 'NA', 'EU', 'JP'] } df = pd.DataFrame(data=d) target_columns = df.columns.tolist() hash_size_per_col = 6 # Create a list of transformer tuples: (name, transformer, target_column) transformer_list = [] for col in target_columns: hasher = FeatureHasher(n_features=hash_size_per_col, input_type='string') transformer_list.append((f"{col}_hasher", hasher, [col])) # Configure ColumnTransformer to apply all hashers column_transformer = ColumnTransformer( transformers=transformer_list, remainder='passthrough' # Keep original categorical columns; use 'drop' to remove them ) # Fit and transform the data hashed_output = column_transformer.fit_transform(df) # Convert back to a DataFrame (ColumnTransformer returns a sparse matrix by default) # Generate clear column names for the hashed features hashed_col_names = [] for name, _, cols in transformer_list: hashed_col_names.extend([f"{cols[0]}_hash_{i}" for i in range(hash_size_per_col)]) # Combine original column names with hashed column names final_column_names = target_columns + hashed_col_names final_df = pd.DataFrame(hashed_output.toarray(), columns=final_column_names) print(final_df.head())
Key Benefits:
- Integrates directly with sklearn pipelines (easily add a model step afterward)
- Automatically handles column alignment and sparse matrix conversion
- Cleaner code for production-grade workflows
内容的提问来源于stack exchange,提问作者Noor

