You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

TensorFlow Dataset中与Pandas DataFrame.info()等价的数据集结构查看方法是什么?

Is there a TensorFlow Dataset equivalent to Pandas DataFrame.info()?

Great question! Unlike Pandas, which has a handy built-in info() method that gives you a full breakdown of your DataFrame's schema (including non-null counts, data types, and memory usage), TensorFlow Datasets don't have a direct out-of-the-box equivalent. But you can easily replicate a similar schema overview with a few lines of code.

Here's how to do it with a CSV-based TF Dataset (using the Titanic dataset as an example):

First, load your dataset with tf.data.experimental.make_csv_dataset, then inspect the first batch to extract feature and label types:

import tensorflow as tf

titanic_file = "path/to/your/titanic.csv"

titanic = tf.data.experimental.make_csv_dataset(
    titanic_file,
    label_name="survived",
    batch_size=1,  # Grab a single row for easy inspection
    shuffle=False,  # Keep order matching the original CSV
    header=True,
)

# Extract the first batch to inspect features and labels
for row in titanic.take(1):
    features = row[0]  # Features are stored in a dictionary
    label = row[1]
    
    # Print each feature's name and data type
    for feature_name, feature_value in features.items():
        print(f"{feature_name:20s}: {feature_value.dtype}")
    # Print the label's data type
    print(f"label/survived : {label.dtype}")

Output:

sex                 : <dtype: 'string'>
age                 : <dtype: 'float32'>
n_siblings_spouses  : <dtype: 'int32'>
parch               : <dtype: 'int32'>
fare                : <dtype: 'float32'>
class               : <dtype: 'string'>
deck                : <dtype: 'string'>
embark_town         : <dtype: 'string'>
alone               : <dtype: 'string'>
label/survived : <dtype: 'int32'>

Notes:

  • This gives you the core schema details: feature names and their corresponding TensorFlow data types, matching the dtype section of Pandas' info().
  • If you want to replicate more of info()'s functionality (like non-null counts or memory usage), you'll need to add extra logic:
    • For non-null checks: Iterate through more batches and count missing values (using tf.math.is_nan for numerical features or checking for empty strings for string features).
    • For memory usage: Calculate based on each feature's dtype and the total number of samples in your dataset.

内容的提问来源于stack exchange,提问作者mon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 22:07:30