使用TensorFlow开发自编码器:如何将文本结构化数据转为数值格式?
Great question! Working with structured text data instead of images for autoencoders requires targeted preprocessing to turn those string and mixed-type columns into tensors TensorFlow can work with. Let’s break down each column in your dataset and walk through exactly how to convert them step by step.
First, load your whitespace-separated data into a pandas DataFrame (this makes preprocessing easier):
import pandas as pd # Skip the header separator line when loading df = pd.read_csv('your_data.txt', sep='\s+', skiprows=[1])
Step 1: Categorize Your Columns
First, split your data into two types to handle them appropriately:
- Numerical columns:
orig_p(port number),trans_depth(integer depth) — these are already numeric but need scaling. - Categorical columns:
uid,orig_h(IP address),method,host— these need conversion from text to numerical representations.
Step 2: Process Numerical Columns
Autoencoders perform best with scaled features. Use TensorFlow’s Normalization layer to standardize values (mean 0, standard deviation 1) or normalize to the [0,1] range:
import tensorflow as tf from tensorflow.keras.layers import Normalization numerical_cols = ['orig_p', 'trans_depth'] numerical_data = df[numerical_cols].values # Create and adapt the normalization layer to your data norm_layer = Normalization() norm_layer.adapt(numerical_data) # Apply normalization to get scaled tensors normalized_numerical = norm_layer(numerical_data)
Step 3: Process Categorical Columns
Each categorical column needs a different approach based on how many unique values it has (cardinality):
A. Low-Cardinality Columns (e.g., method)
Columns like method (only POST/GET and maybe a few others) work well with one-hot encoding. This converts each category into a binary vector:
from tensorflow.keras.layers import CategoryEncoding, StringLookup # First map string values to integers method_lookup = StringLookup(output_mode='int') method_lookup.adapt(df['method'].values) method_int = method_lookup(df['method'].values) # Convert integers to one-hot vectors method_onehot = CategoryEncoding( num_tokens=method_lookup.vocabulary_size(), output_mode='one_hot' )(method_int)
B. IP Address (orig_h)
IPs are structured—split them into 4 numerical octets instead of treating them as arbitrary categories. This preserves the logical structure of the IP address:
# Split IP string into 4 octet columns and convert to integers df[['octet1', 'octet2', 'octet3', 'octet4']] = df['orig_h'].str.split('.', expand=True).astype(int) # Normalize the octet values (same as step 2) ip_cols = ['octet1', 'octet2', 'octet3', 'octet4'] ip_data = df[ip_cols].values ip_norm_layer = Normalization() ip_norm_layer.adapt(ip_data) normalized_ip = ip_norm_layer(ip_data)
C. High-Cardinality Columns (uid, host)
For columns with many unique values (like uid, which is likely unique per row), one-hot encoding would create an enormous vector. Instead, use an embedding layer to map each category to a dense, low-dimensional vector:
# Process 'uid' uid_lookup = StringLookup(output_mode='int') uid_lookup.adapt(df['uid'].values) uid_int = uid_lookup(df['uid'].values) # Create embedding layer (adjust output_dim based on your dataset size) uid_embedding = tf.keras.layers.Embedding( input_dim=uid_lookup.vocabulary_size(), output_dim=8 )(uid_int) # Flatten embedding to match other feature shapes uid_embedding_flat = tf.keras.layers.Flatten()(uid_embedding) # Repeat the same process for 'host' host_lookup = StringLookup(output_mode='int') host_lookup.adapt(df['host'].values) host_int = host_lookup(df['host'].values) host_embedding = tf.keras.layers.Embedding( input_dim=host_lookup.vocabulary_size(), output_dim=8 )(host_int) host_embedding_flat = tf.keras.layers.Flatten()(host_embedding)
Step 4: Combine All Processed Features
Once all columns are converted to numerical tensors, concatenate them into a single input tensor for your autoencoder:
combined_input = tf.keras.layers.concatenate([ normalized_numerical, method_onehot, normalized_ip, uid_embedding_flat, host_embedding_flat ])
Step 5: Build an Autoencoder with Integrated Preprocessing
To make your pipeline reproducible (and avoid manual preprocessing during inference), wrap the preprocessing layers directly into your autoencoder model:
# Define input layers for each raw column uid_input = tf.keras.Input(shape=(1,), dtype=tf.string, name='uid') orig_h_input = tf.keras.Input(shape=(1,), dtype=tf.string, name='orig_h') orig_p_input = tf.keras.Input(shape=(1,), dtype=tf.float32, name='orig_p') trans_depth_input = tf.keras.Input(shape=(1,), dtype=tf.float32, name='trans_depth') method_input = tf.keras.Input(shape=(1,), dtype=tf.string, name='method') host_input = tf.keras.Input(shape=(1,), dtype=tf.string, name='host') # Process IP input with a lambda layer to split octets def split_ip(ip): split = tf.strings.split(ip, '.') return tf.strings.to_number(split, out_type=tf.float32) ip_octets = tf.keras.layers.Lambda(split_ip)(orig_h_input) processed_ip = ip_norm_layer(ip_octets) # Process other inputs processed_numerical = norm_layer(tf.keras.layers.concatenate([orig_p_input, trans_depth_input])) processed_method = CategoryEncoding(num_tokens=method_lookup.vocabulary_size(), output_mode='one_hot')(method_lookup(method_input)) processed_uid = tf.keras.layers.Flatten()(tf.keras.layers.Embedding(input_dim=uid_lookup.vocabulary_size(), output_dim=8)(uid_lookup(uid_input))) processed_host = tf.keras.layers.Flatten()(tf.keras.layers.Embedding(input_dim=host_lookup.vocabulary_size(), output_dim=8)(host_lookup(host_input))) # Combine all processed features combined = tf.keras.layers.concatenate([ processed_numerical, processed_method, processed_ip, processed_uid, processed_host ]) # Build encoder encoder = tf.keras.layers.Dense(64, activation='relu')(combined) encoder = tf.keras.layers.Dense(32, activation='relu')(encoder) latent = tf.keras.layers.Dense(16, activation='relu')(encoder) # Build decoder decoder = tf.keras.layers.Dense(32, activation='relu')(latent) decoder = tf.keras.layers.Dense(64, activation='relu')(decoder) output = tf.keras.layers.Dense(combined.shape[1], activation='sigmoid')(decoder) # Final autoencoder model autoencoder = tf.keras.Model( inputs=[uid_input, orig_h_input, orig_p_input, trans_depth_input, method_input, host_input], outputs=output ) autoencoder.compile(optimizer='adam', loss='mse')
Key Tips
- Cardinality Check: Always count unique values in categorical columns—embeddings are non-negotiable for high-cardinality data to avoid the curse of dimensionality.
- Normalization: Never skip scaling numerical features; autoencoders are highly sensitive to feature scale differences.
- Reproducibility: Wrapping preprocessing into the model ensures raw data can be fed directly during inference without extra steps.
内容的提问来源于stack exchange,提问作者CodeDezk

