You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用TensorFlow结合数值与非数值特征预测价格变动百分比?

解决方案

一、模型选择

这是回归任务(预测连续数值型的price_change_percentage),结合你的结构化数据特征(类别型+数值型),推荐以下几种TensorFlow模型:

1. 基于特征列(Feature Columns)的全连接神经网络

这是处理结构化数据最直接的方案,通过TensorFlow的tf.feature_column模块统一处理不同类型的特征,再输入到全连接层做回归预测。适合数据量不大、特征复杂度中等的场景。

2. 宽深模型(Wide & Deep Model)

如果需要同时捕捉特征的线性关联(Wide部分)和复杂非线性模式(Deep部分),可以用宽深模型。Wide部分处理类别特征的交叉组合,Deep部分学习特征的深层表示,适合兼顾记忆性和泛化性的结构化数据任务。

3. TensorFlow决策森林(TensorFlow Decision Forests)

若偏好模型的可解释性,TF-DF是不错的选择——它原生支持类别特征,无需手动编码,训练速度快且效果稳定,非常适配结构化数据回归任务。


二、非数值特征的转换方法

product_name和brand属于类别型特征,需转换为模型可识别的数值格式,具体方法如下:

1. 独热编码(One-Hot Encoding)

适合类别数量较少的特征(比如你的brand仅X、Y两类)。将每个类别映射为二进制向量,例如:

  • brand=X → [1, 0]
  • brand=Y → [0, 1]

TensorFlow实现示例:

import tensorflow as tf

# 处理brand特征
brand_vocab = ['X', 'Y']
brand_cat_col = tf.feature_column.categorical_column_with_vocabulary_list(
    key='brand', vocabulary_list=brand_vocab
)
brand_one_hot = tf.feature_column.indicator_column(brand_cat_col)

2. 嵌入编码(Embedding)

如果product_name的类别数量较多(比如后续会新增大量不同玩具名称),独热编码会导致特征维度爆炸,此时推荐嵌入编码。它将高维类别特征映射到低维稠密向量,模型可学习到类别间的关联信息。

TensorFlow实现示例:

# 处理product_name特征(用哈希桶适配大量类别)
product_name_hash_col = tf.feature_column.categorical_column_with_hash_bucket(
    key='product_name', hash_bucket_size=100  # 根据实际类别数量调整
)
product_name_embedding = tf.feature_column.embedding_column(
    categorical_column=product_name_hash_col, dimension=8  # 嵌入维度按需调整
)

3. 注意事项

  • 标签编码(Label Encoding)不适合此类无序类别特征,模型会错误认为类别间存在顺序关系(比如Toy X=0、Toy Y=1,模型可能解读为0<1有意义)。
  • release_year属于数值特征,可直接作为数值列输入:
release_year_num_col = tf.feature_column.numeric_column(key='release_year')

三、简单模型实现示例

整合上述特征,构建基础全连接回归模型:

# 整合所有特征列
feature_columns = [brand_one_hot, product_name_embedding, release_year_num_col]

# 构建输入层
input_layer = tf.keras.layers.DenseFeatures(feature_columns)

# 构建模型
model = tf.keras.Sequential([
    input_layer,
    tf.keras.layers.Dense(64, activation='relu'),
    tf.keras.layers.Dense(32, activation='relu'),
    tf.keras.layers.Dense(1)  # 回归任务输出层无需激活函数
])

# 编译模型
model.compile(optimizer='adam', loss='mean_squared_error', metrics=['mae'])

内容的提问来源于stack exchange,提问作者Shawn Hunter

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 07:20:38