You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将数值数据转回类别标签,并基于原始类别数据完成预测?

类别编码反转与模型预测解决方案

问题核心

你通过pandas.Series.astype('category').cat.codes将UCI汽车数据集的所有列转为数值编码并训练了决策树模型,但无法直接输入包含原始类别(如'audi')的混合数据进行预测,同时需要实现编码与原始类别之间的双向转换。


解决方案步骤

1. 单独保存类别列的编码映射

不要一次性批量转换所有列,而是为每个类别列单独创建类别→编码和编码→类别的映射字典,这是实现双向转换的核心。

import pandas as pd
import numpy as np
from sklearn.model_selection import train_test_split
from sklearn.tree import DecisionTreeClassifier
from sklearn.metrics import accuracy_score

# 定义表头并读取数据
headers = ["symboling", "normalized_losses", "make", "fuel_type", "aspiration",
           "num_doors", "body_style", "drive_wheels", "engine_location",
           "wheel_base", "length", "width", "height", "curb_weight",
           "engine_type", "num_cylinders", "engine_size", "fuel_system",
           "bore", "stroke", "compression_ratio", "horsepower", "peak_rpm",
           "city_mpg", "highway_mpg", "price"]

df = pd.read_csv("https://archive.ics.uci.edu/ml/machine-learning-databases/autos/imports-85.data",
                  header=None, names=headers, na_values="?" )

# 处理缺失值(必须步骤,避免编码出错)
df = df.dropna()

# 手动指定类别列(根据数据业务逻辑区分)
category_cols = ["make", "fuel_type", "aspiration", "num_doors", "body_style", 
                 "drive_wheels", "engine_location", "engine_type", "num_cylinders", "fuel_system"]

# 构建编码映射与反转映射
code_maps = {}  # 原始类别 → 编码值
reverse_code_maps = {}  # 编码值 → 原始类别

for col in category_cols:
    df[col] = df[col].astype('category')
    code_maps[col] = dict(zip(df[col].cat.categories, df[col].cat.codes))
    reverse_code_maps[col] = dict(zip(df[col].cat.codes, df[col].cat.categories))

# 构建训练数据集:数值列保留原值,类别列转编码
df_fin = df.copy()
for col in category_cols:
    df_fin[col] = df_fin[col].cat.codes

# 准备特征与目标变量,划分数据集并训练模型
X = df_fin.drop('city_mpg', axis=1)
y = df_fin['city_mpg']

X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
clf = DecisionTreeClassifier()
clf.fit(X_train, y_train)

# 模型评估
y_pred = clf.predict(X_test)
print(f"模型准确率:{accuracy_score(y_test, y_pred):.2f}")

2. 处理混合输入数据并预测

将新输入中的类别字段通过之前保存的code_maps转换为编码值,数值字段直接保留,再传入模型预测;若需要将预测结果(编码值)转回原始类别/数值,使用反转映射。

# 示例混合输入:包含数值和原始类别
new_input = [2,164,'audi','gas','std','four','sedan','fwd','front',99.8,176.6,66.2,54.3,2337,'ohc','four',109,'mpfi',3.19,3.4,10,102,5500,30,13950]

# 转换为DataFrame以便按列处理
input_df = pd.DataFrame([new_input], columns=X.columns)

# 将输入中的类别字段转为编码值
for col in category_cols:
    input_df[col] = input_df[col].map(code_maps[col])

# 执行预测
predicted_code = clf.predict(input_df)[0]

# (可选)将预测的编码值转回原始city_mpg数值
# 先构建city_mpg的反转映射
city_mpg_cat = df['city_mpg'].astype('category')
city_mpg_reverse_map = dict(zip(city_mpg_cat.codes, city_mpg_cat.categories))
predicted_city_mpg = city_mpg_reverse_map[predicted_code]

print(f"预测的城市油耗:{predicted_city_mpg}")

3. 关键注意事项

  • 区分数值列与类别列:无需将所有列转为类别编码,数值型字段(如wheel_base、city_mpg)直接保留原值,避免丢失数值连续性信息。
  • 必须保存映射字典:训练后不能丢弃code_maps和reverse_code_maps,否则无法实现新输入的编码转换和预测结果的反转。
  • 缺失值处理:原始数据中的缺失值会导致编码失败,训练前必须通过删除、填充等方式处理。

内容的提问来源于stack exchange,提问作者ASH

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 20:35:16