You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PyTorch中拆分训练测试集后,如何查看其中的具体数据?

查看random_split拆分后的训练/测试集数据

方法1:通过Subset索引从原始DataFrame提取(最便捷)

random_split返回的Subset对象自带indices属性,存储了该子集在原始Dataset中的索引位置。直接用这些索引从最初的pandas DataFrame提取数据,能完全保留原始数据格式:

# 获取训练集在原始DataFrame中的索引并提取数据
train_indices = train_set.indices
train_df = df.iloc[train_indices]
print("训练集数据:")
print(train_df)

# 获取测试集在原始DataFrame中的索引并提取数据
test_indices = test_set.indices
test_df = df.iloc[test_indices]
print("\n测试集数据:")
print(test_df)

方法2:遍历Subset转成DataFrame

如果你的Dataset是TensorDataset或自定义Dataset,可遍历子集内的每个样本,整理成字典后转成pandas DataFrame:

# 处理训练集
train_samples = []
for hour_tensor, score_tensor in train_set:
    train_samples.append({
        'hour': hour_tensor.item(),
        'score': score_tensor.item()
    })
train_df = pd.DataFrame(train_samples)
print(train_df.head())

# 处理测试集同理
test_samples = []
for hour_tensor, score_tensor in test_set:
    test_samples.append({
        'hour': hour_tensor.item(),
        'score': score_tensor.item()
    })
test_df = pd.DataFrame(test_samples)
print(test_df.head())

方法3:查看单个或少量样本

如果只是快速验证几个样本,直接通过索引访问Subset对象即可:

# 查看训练集第3个样本
hour, score = train_set[3]
print(f"训练集第3个样本:hour={hour.item()}, score={score.item()}")

# 查看测试集前5个样本
print("\n测试集前5个样本:")
for i in range(5):
    hour, score = test_set[i]
    print(f"样本{i}: hour={hour.item()}, score={score.item()}")

内容的提问来源于stack exchange,提问作者Yeongkyu Lee

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 12:50:24