如何在Dask DataFrame中使用Scikit-learn的FeatureHasher?
Using FeatureHasher with Dask DataFrames Correctly
Got it, let's work through how to properly apply FeatureHasher to your Dask DataFrame—this is a common snag when dealing with high-cardinality categorical features, so I'll break it down clearly.
First, let's address why you're getting such a short array:
- If you tried using scikit-learn's vanilla
FeatureHasherdirectly on the Dask DataFrame, it only processes the first partition (not the entire dataset), hence the tiny output. - The
n_featuresparameter also plays a huge role: its default value (2**18 = 262,144) might be too small for your high-cardinality features, leading to a smaller-than-expected feature space.
Below are two reliable approaches, with the first being the most straightforward for Dask workflows:
Approach 1: Use Dask-ML's Native FeatureHasher (Recommended)
Dask-ML provides a FeatureHasher that's built to work seamlessly with Dask's distributed data structures, so you don't have to handle partition mapping manually.
Step-by-Step Code
import dask.dataframe as dd from dask_ml.feature_extraction import FeatureHasher # Load your dataset (adjust the read method to match your data source) df = dd.read_csv('your_data.csv') # Define the categorical features you want to hash categorical_cols = ['gender', 'category', 'subcategory', 'item_brand', 'item_NWT'] # Initialize the FeatureHasher with a suitable n_features value # Pick a power of 2 (efficient for hashing) that's large enough to minimize collisions # For high-cardinality features, 2**20 (1,048,576) is a good starting point hasher = FeatureHasher( n_features=2**20, input_type='string' # Use 'string' since your features are text-based ) # Generate hashed features hashed_features = hasher.fit_transform(df[categorical_cols]) # Optional: Convert hashed features to a Dask DataFrame to join with your original data hashed_df = hashed_features.to_dask_dataframe() final_df = df.join(hashed_df)
Approach 2: Scikit-Learn FeatureHasher + Dask map_partitions
If you prefer sticking to scikit-learn's implementation, you can use Dask's map_partitions to apply the hasher across all dataset partitions.
Step-by-Step Code
import dask.dataframe as dd from sklearn.feature_extraction import FeatureHasher def hash_single_partition(partition): # Initialize the hasher inside the function (safe for distributed execution) hasher = FeatureHasher(n_features=2**20, input_type='dict') # Convert each row in the partition to a dictionary (required by scikit-learn's hasher) row_dicts = partition.to_dict('records') return hasher.transform(row_dicts) # Load your data df = dd.read_csv('your_data.csv') categorical_cols = ['gender', 'category', 'subcategory', 'item_brand', 'item_NWT'] # Apply the hasher to all partitions # Specify meta to tell Dask the output data type hashed_features = df[categorical_cols].map_partitions( hash_single_partition, meta=('feature', 'float64') ) # Optional: Convert to DataFrame and join with original data hashed_df = hashed_features.to_dask_dataframe() final_df = df.join(hashed_df)
Key Notes to Avoid Issues
- Adjust
n_features: Start with a value like 220 or 222 if your features have extremely high cardinality. Too small a value increases hash collisions, which hurts model performance; too large wastes compute resources. - Input Type: Use
input_type='string'for Dask-ML if your columns are raw text. For scikit-learn,input_type='dict'works since we're converting rows to dictionaries. - Validate Output: Run
hashed_features.head().compute()to check the feature dimensions—they should match then_featuresyou set.
内容的提问来源于stack exchange,提问作者Jorge
相关产品推荐
相关产品推荐

