如何仅对DataFrame中的分类数据进行编码?
Hey there! Let's walk through how to encode the categorical data in your DataFrame, based on the sample data you shared. First, let's nail down which columns are actually categorical—looking at your fields and sample rows, these are the non-numeric (or discrete, unordered) columns we need to handle:
- Median House Value: Takes values like
FandH(likely categorical tiers) - MakeDescr: Car brands (e.g.,
HYUNDAI,JEEP,VOLKSWAGEN) - SeriesDescr: Car model series (e.g.,
Sonata-4 Cyl.,Grand Cherokee-V6)
All other columns (like Income, Percentage Asian, Msrp) are numeric and don't need categorical encoding. Below are the most common methods to encode these categorical columns, with code examples tailored to your data.
1. First, Build the Sample DataFrame
Let's start by recreating your sample data in pandas so we can test the encoding:
import pandas as pd # Sample data from your input data = { 'Income': [90000, 125000, 90000, 125000], 'Length of Residence': [15.0, 7.0, 8.0, 8.0], 'Median House Value': ['F', 'H', 'F', 'F'], 'Number of Vehicles': [4, 1, 1, 3], 'Percentage Asian': [1, 11, 1, 1], 'Percentage Black': [1, 1, 1, 1], 'Percentage English Speaking': [71, 91, 71, 86], 'Percentage Hispanic': [6, 1, 6, 6], 'Percentage White': [81, 81, 86, 86], 'MakeDescr': ['HYUNDAI', 'JEEP', 'JEEP', 'VOLKSWAGEN'], 'SeriesDescr': ['Sonata-4 Cyl.', 'Grand Cherokee-V6', 'Liberty', 'Passat-V6'], 'Msrp': [19395.0, 29135.0, 20700.0, 28750.0] } df = pd.DataFrame(data)
2. Label Encoding (For Ordinal Categorical Data)
Use this if your categorical column has an inherent order (e.g., Median House Value where F might stand for "Fair" and H for "High"—a clear hierarchy). Label encoding assigns a unique integer to each category:
from sklearn.preprocessing import LabelEncoder # Initialize the encoder le = LabelEncoder() # Apply encoding to Median House Value df['Median House Value_Encoded'] = le.fit_transform(df['Median House Value']) # Check the result print(df[['Median House Value', 'Median House Value_Encoded']])
Output will look like:
Median House Value Median House Value_Encoded 0 F 0 1 H 1 2 F 0 3 F 0
3. One-Hot Encoding (For Nominal Categorical Data)
Use this for unordered categories (like car brands or models—there's no "rank" between HYUNDAI and JEEP). This creates a new binary column for each category (excluding one to avoid multicollinearity):
Option A: Pandas get_dummies (Quick & Simple)
# One-hot encode MakeDescr and SeriesDescr df_encoded = pd.get_dummies(df, columns=['MakeDescr', 'SeriesDescr'], drop_first=True) # View the new columns print(df_encoded.columns.tolist())
Option B: Scikit-Learn OneHotEncoder (Better for Pipelines)
If you're planning to build a machine learning pipeline, use this method—it integrates seamlessly with other preprocessing steps:
from sklearn.preprocessing import OneHotEncoder from sklearn.compose import ColumnTransformer # Define which columns to encode categorical_cols = ['Median House Value', 'MakeDescr', 'SeriesDescr'] # Build a column transformer: encode categoricals, leave numeric columns as-is preprocessor = ColumnTransformer( transformers=[ ('cat_encoder', OneHotEncoder(sparse_output=False, drop='first'), categorical_cols) ], remainder='passthrough' ) # Apply the transformation encoded_data = preprocessor.fit_transform(df) # Convert back to DataFrame for readability encoded_df = pd.DataFrame( encoded_data, columns=preprocessor.get_feature_names_out() ) print(encoded_df.head())
Key Notes to Remember
- Ordinal vs Nominal: Always pick the right encoding method based on whether your category has an order (label encode) or not (one-hot encode).
- High Cardinality: If columns like
SeriesDescrhave hundreds of unique values, one-hot encoding will bloat your feature space. In that case, consider target encoding (using the target variable's mean per category) or frequency encoding—but be cautious of overfitting! - Missing Values: Before encoding, check for missing values in categorical columns. You can fill them with the most frequent category or a placeholder like "Unknown".
内容的提问来源于stack exchange,提问作者Ajay Kumar

