You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何无需转换为数值,使用字符串特征与标签训练决策树分类器?

Awesome question—you don't need to manually map those string features and labels to numbers anymore! Scikit-learn has built-in tools that handle this automatically, making your code cleaner and less error-prone. Here's how to implement it:

Step 1: Understand the Tools We'll Use

  • ColumnTransformer: Lets us apply different preprocessing steps to specific columns (critical here, since we have one numeric column and one categorical string column).
  • OrdinalEncoder: Converts categorical string values (like "smooth"/"bumpy") into numerical values consistently.
  • Pipeline: Combines our preprocessing steps and classifier into a single, reusable object—this avoids manual data transformation and helps prevent data leakage.

Full Working Code

from sklearn.tree import DecisionTreeClassifier
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import OrdinalEncoder
from sklearn.pipeline import Pipeline

# Your original string-containing data (no manual conversion needed!)
features = [[140,"smooth"],[130,"smooth"],[150,"bumpy"],[170,"bumpy"]]
labels = ["apple","apple","orange","orange"]

# Set up preprocessing: encode the 2nd column (index 1), leave 1st column as-is
preprocessor = ColumnTransformer(
    transformers=[
        ('categorical_encoder', OrdinalEncoder(), [1])
    ],
    remainder='passthrough'
)

# Create a pipeline that preprocesses data then trains the classifier
model = Pipeline(steps=[
    ('preprocessor', preprocessor),
    ('classifier', DecisionTreeClassifier())
])

# Train the model directly on your original data
model.fit(features, labels)

# Make a prediction using string values just like your input data
prediction = model.predict([[150, "bumpy"]])
print(prediction)  # Output: ['orange']

Key Notes

  1. No Manual Label Conversion: Scikit-learn's classifiers (including DecisionTreeClassifier) accept string labels directly—you don't need to convert "apple"/"orange" to numbers at all.
  2. Consistent Encoding: OrdinalEncoder handles the string-to-number mapping for features automatically, and the pipeline ensures this same mapping is used for predictions (so you never have to remember if "smooth" was 0 or 1!).
  3. Alternative: Manual Preprocessing (If You Prefer)
    If you don't want to use a pipeline, you can transform the features first then train the classifier:
    # Manually transform features
    transformed_features = preprocessor.fit_transform(features)
    
    # Train classifier on transformed features
    classifier = DecisionTreeClassifier()
    classifier.fit(transformed_features, labels)
    
    # Don't forget to transform new prediction data too!
    prediction = classifier.predict(preprocessor.transform([[150, "bumpy"]]))
    print(prediction)
    

内容的提问来源于stack exchange,提问作者Ashish.k

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 07:16:53