You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用TensorFlow开发自编码器:如何将文本结构化数据转为数值格式?

Converting Structured Text Data to TensorFlow-Compatible Numerical Format for Autoencoders

Great question! Working with structured text data instead of images for autoencoders requires targeted preprocessing to turn those string and mixed-type columns into tensors TensorFlow can work with. Let’s break down each column in your dataset and walk through exactly how to convert them step by step.

First, load your whitespace-separated data into a pandas DataFrame (this makes preprocessing easier):

import pandas as pd
# Skip the header separator line when loading
df = pd.read_csv('your_data.txt', sep='\s+', skiprows=[1])

Step 1: Categorize Your Columns

First, split your data into two types to handle them appropriately:

  • Numerical columns: orig_p (port number), trans_depth (integer depth) — these are already numeric but need scaling.
  • Categorical columns: uid, orig_h (IP address), method, host — these need conversion from text to numerical representations.

Step 2: Process Numerical Columns

Autoencoders perform best with scaled features. Use TensorFlow’s Normalization layer to standardize values (mean 0, standard deviation 1) or normalize to the [0,1] range:

import tensorflow as tf
from tensorflow.keras.layers import Normalization

numerical_cols = ['orig_p', 'trans_depth']
numerical_data = df[numerical_cols].values

# Create and adapt the normalization layer to your data
norm_layer = Normalization()
norm_layer.adapt(numerical_data)

# Apply normalization to get scaled tensors
normalized_numerical = norm_layer(numerical_data)

Step 3: Process Categorical Columns

Each categorical column needs a different approach based on how many unique values it has (cardinality):

A. Low-Cardinality Columns (e.g., method)

Columns like method (only POST/GET and maybe a few others) work well with one-hot encoding. This converts each category into a binary vector:

from tensorflow.keras.layers import CategoryEncoding, StringLookup

# First map string values to integers
method_lookup = StringLookup(output_mode='int')
method_lookup.adapt(df['method'].values)
method_int = method_lookup(df['method'].values)

# Convert integers to one-hot vectors
method_onehot = CategoryEncoding(
    num_tokens=method_lookup.vocabulary_size(), 
    output_mode='one_hot'
)(method_int)

B. IP Address (orig_h)

IPs are structured—split them into 4 numerical octets instead of treating them as arbitrary categories. This preserves the logical structure of the IP address:

# Split IP string into 4 octet columns and convert to integers
df[['octet1', 'octet2', 'octet3', 'octet4']] = df['orig_h'].str.split('.', expand=True).astype(int)

# Normalize the octet values (same as step 2)
ip_cols = ['octet1', 'octet2', 'octet3', 'octet4']
ip_data = df[ip_cols].values
ip_norm_layer = Normalization()
ip_norm_layer.adapt(ip_data)
normalized_ip = ip_norm_layer(ip_data)

C. High-Cardinality Columns (uid, host)

For columns with many unique values (like uid, which is likely unique per row), one-hot encoding would create an enormous vector. Instead, use an embedding layer to map each category to a dense, low-dimensional vector:

# Process 'uid'
uid_lookup = StringLookup(output_mode='int')
uid_lookup.adapt(df['uid'].values)
uid_int = uid_lookup(df['uid'].values)

# Create embedding layer (adjust output_dim based on your dataset size)
uid_embedding = tf.keras.layers.Embedding(
    input_dim=uid_lookup.vocabulary_size(), 
    output_dim=8
)(uid_int)
# Flatten embedding to match other feature shapes
uid_embedding_flat = tf.keras.layers.Flatten()(uid_embedding)

# Repeat the same process for 'host'
host_lookup = StringLookup(output_mode='int')
host_lookup.adapt(df['host'].values)
host_int = host_lookup(df['host'].values)
host_embedding = tf.keras.layers.Embedding(
    input_dim=host_lookup.vocabulary_size(), 
    output_dim=8
)(host_int)
host_embedding_flat = tf.keras.layers.Flatten()(host_embedding)

Step 4: Combine All Processed Features

Once all columns are converted to numerical tensors, concatenate them into a single input tensor for your autoencoder:

combined_input = tf.keras.layers.concatenate([
    normalized_numerical,
    method_onehot,
    normalized_ip,
    uid_embedding_flat,
    host_embedding_flat
])

Step 5: Build an Autoencoder with Integrated Preprocessing

To make your pipeline reproducible (and avoid manual preprocessing during inference), wrap the preprocessing layers directly into your autoencoder model:

# Define input layers for each raw column
uid_input = tf.keras.Input(shape=(1,), dtype=tf.string, name='uid')
orig_h_input = tf.keras.Input(shape=(1,), dtype=tf.string, name='orig_h')
orig_p_input = tf.keras.Input(shape=(1,), dtype=tf.float32, name='orig_p')
trans_depth_input = tf.keras.Input(shape=(1,), dtype=tf.float32, name='trans_depth')
method_input = tf.keras.Input(shape=(1,), dtype=tf.string, name='method')
host_input = tf.keras.Input(shape=(1,), dtype=tf.string, name='host')

# Process IP input with a lambda layer to split octets
def split_ip(ip):
    split = tf.strings.split(ip, '.')
    return tf.strings.to_number(split, out_type=tf.float32)
ip_octets = tf.keras.layers.Lambda(split_ip)(orig_h_input)
processed_ip = ip_norm_layer(ip_octets)

# Process other inputs
processed_numerical = norm_layer(tf.keras.layers.concatenate([orig_p_input, trans_depth_input]))
processed_method = CategoryEncoding(num_tokens=method_lookup.vocabulary_size(), output_mode='one_hot')(method_lookup(method_input))
processed_uid = tf.keras.layers.Flatten()(tf.keras.layers.Embedding(input_dim=uid_lookup.vocabulary_size(), output_dim=8)(uid_lookup(uid_input)))
processed_host = tf.keras.layers.Flatten()(tf.keras.layers.Embedding(input_dim=host_lookup.vocabulary_size(), output_dim=8)(host_lookup(host_input)))

# Combine all processed features
combined = tf.keras.layers.concatenate([
    processed_numerical, processed_method, processed_ip, processed_uid, processed_host
])

# Build encoder
encoder = tf.keras.layers.Dense(64, activation='relu')(combined)
encoder = tf.keras.layers.Dense(32, activation='relu')(encoder)
latent = tf.keras.layers.Dense(16, activation='relu')(encoder)

# Build decoder
decoder = tf.keras.layers.Dense(32, activation='relu')(latent)
decoder = tf.keras.layers.Dense(64, activation='relu')(decoder)
output = tf.keras.layers.Dense(combined.shape[1], activation='sigmoid')(decoder)

# Final autoencoder model
autoencoder = tf.keras.Model(
    inputs=[uid_input, orig_h_input, orig_p_input, trans_depth_input, method_input, host_input],
    outputs=output
)

autoencoder.compile(optimizer='adam', loss='mse')

Key Tips

  • Cardinality Check: Always count unique values in categorical columns—embeddings are non-negotiable for high-cardinality data to avoid the curse of dimensionality.
  • Normalization: Never skip scaling numerical features; autoencoders are highly sensitive to feature scale differences.
  • Reproducibility: Wrapping preprocessing into the model ensures raw data can be fed directly during inference without extra steps.

内容的提问来源于stack exchange,提问作者CodeDezk

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 09:19:37