You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

机器学习中分类属性数值化:numpy.unique()与get_dummies()孰优?

Which Categorical Encoding Method is Better in Machine Learning: numpy.unique() vs pd.get_dummies()?

Great question! This is one of those classic "it depends" scenarios in ML preprocessing—there's no universal "better" choice, but the right pick hinges on your model type and the meaning behind your categorical data. Let's break it down clearly:

First, Let's Recap What Each Does

  • numpy.unique() (often called Label Encoding when used for this purpose): Assigns a unique integer (0, 1, 2, ..., K-1) to each category, turning your single categorical column into one numeric column. For example, ["red", "blue", "red"] becomes [0, 1, 0].
  • pd.get_dummies() (One-Hot Encoding): Creates K separate binary columns (one per category), where each column has 1 if the row belongs to that category, 0 otherwise. Using the same example, you'd get two columns: red [1,0,1] and blue [0,1,0].

When to Prefer numpy.unique() (Label Encoding)

  • You're using tree-based models (Decision Trees, Random Forests, XGBoost, LightGBM): Tree models don't care about the "magnitude" of numeric values—they split based on thresholds. Label encoding won't introduce artificial order biases here, and it keeps your feature space small, which is efficient especially with high-cardinality data (categories with dozens/hundreds of distinct values).
  • Your categorical variable is ordinal (has a natural order): If your categories have a meaningful hierarchy (e.g., ["poor", "average", "good"], ["small", "medium", "large"]), label encoding preserves that order information. A model can learn that 2 ("good") is better than 0 ("poor"), which is useful.
  • You need to save memory/compute: Label encoding adds only one column, whereas one-hot encoding adds K columns—this is a huge win if you have a lot of categories.

Caveat with Label Encoding

Never use it for nominal (unordered) variables with models that assume linear relationships (like Linear Regression, Logistic Regression, SVMs, or Neural Networks). For example, encoding ["red", "blue", "green"] as 0,1,2 will make the model think "green" (2) is "twice as much" as "red" (0)—which is totally meaningless and will lead to bad predictions.

When to Prefer pd.get_dummies() (One-Hot Encoding)

  • You're using linear models, SVMs, or neural networks: These models treat numeric features as continuous, so one-hot encoding ensures no artificial order is imposed. Each category is treated as an independent binary feature, which aligns with the models' assumptions.
  • Your categorical variable is nominal (unordered): For variables like ["cat", "dog", "bird"], ["New York", "London", "Tokyo"], one-hot encoding is the safe choice—no false relationships are introduced.
  • You have low-cardinality data: If your category has only 2-5 distinct values, one-hot encoding won't bloat your feature space too much, and it's easy for models to interpret.

Caveat with One-Hot Encoding

Avoid it for high-cardinality variables (e.g., user IDs, product SKUs with thousands of distinct values). This will cause the "curse of dimensionality"—your feature space explodes, leading to slower training, increased memory usage, and higher risk of overfitting. In these cases, you'd want to use other methods like target encoding or embeddings instead.

Quick Decision Cheat Sheet

  1. Ordinal category + any model → Label Encoding (numpy.unique())
  2. Nominal category + tree-based model → Either works, but Label Encoding is more efficient
  3. Nominal category + linear/NN/SVM → One-Hot Encoding (pd.get_dummies())
  4. High-cardinality nominal category → Avoid both, use target encoding/embeddings

内容的提问来源于stack exchange,提问作者Ritvik Khanna

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 09:51:41