TensorFlow Dataset中与Pandas DataFrame.info()等价的数据集结构查看方法是什么?
Is there a TensorFlow Dataset equivalent to Pandas DataFrame.info()?
Great question! Unlike Pandas, which has a handy built-in info() method that gives you a full breakdown of your DataFrame's schema (including non-null counts, data types, and memory usage), TensorFlow Datasets don't have a direct out-of-the-box equivalent. But you can easily replicate a similar schema overview with a few lines of code.
Here's how to do it with a CSV-based TF Dataset (using the Titanic dataset as an example):
First, load your dataset with tf.data.experimental.make_csv_dataset, then inspect the first batch to extract feature and label types:
import tensorflow as tf titanic_file = "path/to/your/titanic.csv" titanic = tf.data.experimental.make_csv_dataset( titanic_file, label_name="survived", batch_size=1, # Grab a single row for easy inspection shuffle=False, # Keep order matching the original CSV header=True, ) # Extract the first batch to inspect features and labels for row in titanic.take(1): features = row[0] # Features are stored in a dictionary label = row[1] # Print each feature's name and data type for feature_name, feature_value in features.items(): print(f"{feature_name:20s}: {feature_value.dtype}") # Print the label's data type print(f"label/survived : {label.dtype}")
Output:
sex : <dtype: 'string'> age : <dtype: 'float32'> n_siblings_spouses : <dtype: 'int32'> parch : <dtype: 'int32'> fare : <dtype: 'float32'> class : <dtype: 'string'> deck : <dtype: 'string'> embark_town : <dtype: 'string'> alone : <dtype: 'string'> label/survived : <dtype: 'int32'>
Notes:
- This gives you the core schema details: feature names and their corresponding TensorFlow data types, matching the dtype section of Pandas'
info(). - If you want to replicate more of
info()'s functionality (like non-null counts or memory usage), you'll need to add extra logic:- For non-null checks: Iterate through more batches and count missing values (using
tf.math.is_nanfor numerical features or checking for empty strings for string features). - For memory usage: Calculate based on each feature's dtype and the total number of samples in your dataset.
- For non-null checks: Iterate through more batches and count missing values (using
内容的提问来源于stack exchange,提问作者mon
相关产品推荐
相关产品推荐

