如何将整数列表及DataFrame中的对象类型列表列转换为TensorFlow Dataset
1. Converting an Integer List to TensorFlow Dataset
Turning a plain integer list into a TensorFlow Dataset is super straightforward with tf.data.Dataset.from_tensor_slices(). This method takes your list (which gets converted to a tensor under the hood) and splits it into individual elements for the dataset.
Here’s a concrete example:
import tensorflow as tf # Your integer list int_list = [2, 4, 6, 8, 10] # Convert to TensorFlow Dataset dataset = tf.data.Dataset.from_tensor_slices(int_list) # Test by iterating through the dataset for item in dataset: print(item.numpy())
This will print each integer as a scalar tensor. If your list is multi-dimensional (like a list of lists), the same method works—it’ll split along the first dimension, making each sublist a dataset element.
2. Converting a Pandas DataFrame Column of Variable-Length Lists to TensorFlow Dataset
Since your DataFrame column has variable-length lists stored as object type, we need a way to handle non-uniform sequence lengths. TensorFlow’s RaggedTensor was built exactly for this scenario—it lets you work with sequences of different lengths without padding.
Here’s how to do it step by step:
import pandas as pd import tensorflow as tf # Sample DataFrame matching your structure df = pd.DataFrame({ 'values': [[0, 2], [0], [5, 1, 9], [7], [1, 3, 5, 7]] }) # Convert the 'values' column to a RaggedTensor ragged_tensor = tf.ragged.constant(df['values'].tolist()) # Create the Dataset from the RaggedTensor dataset = tf.data.Dataset.from_tensor_slices(ragged_tensor) # Verify the output for elem in dataset: print(elem.numpy())
This keeps the original varying lengths of each list intact. If you later need uniform-length sequences (for example, to feed into a neural network), you can use the padded_batch() method on the dataset to add padding and standardize shapes.
内容的提问来源于stack exchange,提问作者Mohsen Mahmoodzadeh

