You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

多分类特征Feature Hashing批量实现:替代手动逐个处理的高效方案

Efficient Feature Hashing for Multiple Categorical Columns

Great question! When you’re dealing with dozens of categorical features (like 40 in your scenario), manually setting up a separate FeatureHasher for each column is tedious and error-prone. Let’s walk through two scalable, clean approaches to solve this problem—plus fix a small typo in your original code first.

Quick Fix for Your Original Code

You initialized fh1 and fh2 but tried to call fh.fit_transform() (where fh isn’t defined) — that’s a typo that would throw an error. The correct calls should use fh1 for Genre and fh2 for Publisher.


Approach 1: Batch Processing with a Loop

This method is straightforward and easy to customize. We’ll loop through all target categorical columns, apply FeatureHasher to each, and combine the results with the original data.

import pandas as pd
from sklearn.feature_extraction import FeatureHasher

# Sample data (scaled to mimic your 40-feature use case)
d = {
    'Genre': ['Platform', 'Racing','Sports','Roleplaying','Puzzle','Platform'],
    'Publisher': ['Nintendo', 'Noir','Laura','John','John','Noir'],
    'Developer': ['DevA', 'DevB', 'DevC', 'DevA', 'DevB', 'DevC'],
    'Region': ['NA', 'EU', 'JP', 'NA', 'EU', 'JP']
    # Add 36 more categorical columns here
}
df = pd.DataFrame(data=d)

# Configuration: Define which columns to hash and hash size per column
target_columns = df.columns.tolist()  # Use a subset like ['Genre', 'Publisher', ...] if needed
hash_size_per_col = 6

# Start with the original categorical columns
processed_dfs = [df[target_columns]]

# Loop through each column to generate hashed features
for col in target_columns:
    # Initialize hasher for the column
    hasher = FeatureHasher(n_features=hash_size_per_col, input_type='string')
    # Transform and convert to array/DataFrame
    hashed_array = hasher.fit_transform(df[col]).toarray()
    # Name columns clearly to avoid duplicate "0/1/2..." labels
    hashed_df = pd.DataFrame(hashed_array, columns=[f"{col}_hash_{i}" for i in range(hash_size_per_col)])
    processed_dfs.append(hashed_df)

# Combine all into one final DataFrame
final_df = pd.concat(processed_dfs, axis=1)
print(final_df.head())

Key Benefits:

  • Simple to understand and modify
  • Gives you full control over column naming (critical for debugging 40+ features)
  • No extra sklearn dependencies beyond FeatureHasher

If you plan to integrate this into a machine learning workflow, sklearn.compose.ColumnTransformer is the more elegant solution. It lets you define all hashers in one place and seamlessly combines results.

import pandas as pd
from sklearn.feature_extraction import FeatureHasher
from sklearn.compose import ColumnTransformer

# Same sample data as above
d = {
    'Genre': ['Platform', 'Racing','Sports','Roleplaying','Puzzle','Platform'],
    'Publisher': ['Nintendo', 'Noir','Laura','John','John','Noir'],
    'Developer': ['DevA', 'DevB', 'DevC', 'DevA', 'DevB', 'DevC'],
    'Region': ['NA', 'EU', 'JP', 'NA', 'EU', 'JP']
}
df = pd.DataFrame(data=d)

target_columns = df.columns.tolist()
hash_size_per_col = 6

# Create a list of transformer tuples: (name, transformer, target_column)
transformer_list = []
for col in target_columns:
    hasher = FeatureHasher(n_features=hash_size_per_col, input_type='string')
    transformer_list.append((f"{col}_hasher", hasher, [col]))

# Configure ColumnTransformer to apply all hashers
column_transformer = ColumnTransformer(
    transformers=transformer_list,
    remainder='passthrough'  # Keep original categorical columns; use 'drop' to remove them
)

# Fit and transform the data
hashed_output = column_transformer.fit_transform(df)

# Convert back to a DataFrame (ColumnTransformer returns a sparse matrix by default)
# Generate clear column names for the hashed features
hashed_col_names = []
for name, _, cols in transformer_list:
    hashed_col_names.extend([f"{cols[0]}_hash_{i}" for i in range(hash_size_per_col)])
# Combine original column names with hashed column names
final_column_names = target_columns + hashed_col_names

final_df = pd.DataFrame(hashed_output.toarray(), columns=final_column_names)
print(final_df.head())

Key Benefits:

  • Integrates directly with sklearn pipelines (easily add a model step afterward)
  • Automatically handles column alignment and sparse matrix conversion
  • Cleaner code for production-grade workflows

内容的提问来源于stack exchange,提问作者Noor

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 09:06:07