You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

能否在Amazon SageMaker自定义处理作业中实现自包含数据可视化?

Solution: Self-Contained Visualization in Amazon SageMaker Processing Jobs

Instead of relying on regex-based metric tracking (which is impractical for large data matrices), you can embed visualization logic directly into your processing job and store outputs in S3 for easy access. Here are the most straightforward, native approaches:

1. Generate Visualizations in the Processing Script & Save to S3

SageMaker Processing jobs automatically sync files in the /opt/ml/output/data/ directory to your specified S3 bucket. You can use standard plotting libraries (Matplotlib, Seaborn, Plotly) to generate pre- and post-transformation visuals, then save them to this output directory.

Example Code Snippet

import pandas as pd
import seaborn as sns
import matplotlib.pyplot as plt
import os
from sklearn.decomposition import PCA

# Load raw input data (from processing job's input channel)
input_dir = "/opt/ml/input/data/raw"
raw_df = pd.read_csv(os.path.join(input_dir, "raw_dataset.csv"))

# --- Pre-processing visualization ---
# Generate correlation heatmap (subset columns if needed for large datasets)
plt.figure(figsize=(14, 10))
sns.heatmap(raw_df.corr(numeric_only=True), annot=False, cmap="viridis")
plt.title("Pre-Processing Feature Correlation")

# Save to output directory (auto-synced to S3)
output_dir = "/opt/ml/output/data/visualizations"
os.makedirs(output_dir, exist_ok=True)
plt.savefig(os.path.join(output_dir, "pre_correlation_heatmap.png"))
plt.close()

# Generate PCA plot for high-dimensional data
pca = PCA(n_components=2)
raw_pca = pca.fit_transform(raw_df.select_dtypes(include='number'))
raw_pca_df = pd.DataFrame({"PC1": raw_pca[:,0], "PC2": raw_pca[:,1]})
plt.figure(figsize=(10, 8))
plt.scatter(raw_pca_df["PC1"], raw_pca_df["PC2"], alpha=0.5)
plt.title("Pre-Processing PCA Projection")
plt.savefig(os.path.join(output_dir, "pre_pca_plot.png"))
plt.close()

# --- Perform your data transformation here ---
transformed_df = raw_df.apply(your_transformation_logic)

# --- Post-processing visualization ---
plt.figure(figsize=(14, 10))
sns.heatmap(transformed_df.corr(numeric_only=True), annot=False, cmap="viridis")
plt.title("Post-Processing Feature Correlation")
plt.savefig(os.path.join(output_dir, "post_correlation_heatmap.png"))
plt.close()

# Post-transformation PCA plot
transformed_pca = pca.fit_transform(transformed_df.select_dtypes(include='number'))
transformed_pca_df = pd.DataFrame({"PC1": transformed_pca[:,0], "PC2": transformed_pca[:,1]})
plt.figure(figsize=(10, 8))
plt.scatter(transformed_pca_df["PC1"], transformed_pca_df["PC2"], alpha=0.5)
plt.title("Post-Processing PCA Projection")
plt.savefig(os.path.join(output_dir, "post_pca_plot.png"))
plt.close()

When defining your processing job, specify the output S3 path in the ProcessingOutput configuration—all files in /opt/ml/output/data/visualizations will be uploaded there once the job completes.

2. View Visualizations in SageMaker Studio

Once the visuals are in S3, you can access them directly in SageMaker Studio:

  • Navigate to your S3 bucket via the Studio file browser, right-click on an image file, and select "Open" to view it.
  • For interactive plots (e.g., Plotly HTML files), open them in Studio's built-in web viewer.

Alternatively, use a notebook cell to display images inline:

from IPython.display import Image
import sagemaker

s3_session = sagemaker.Session().s3_resource
bucket = "your-bucket-name"
prefix = "processing-job-output/visualizations"

# Display pre-processing heatmap
pre_heatmap = s3_session.Object(bucket, f"{prefix}/pre_correlation_heatmap.png").get()["Body"].read()
Image(pre_heatmap)

# Display post-processing PCA plot
post_pca = s3_session.Object(bucket, f"{prefix}/post_pca_plot.png").get()["Body"].read()
Image(post_pca)

3. Optimizations for Large Datasets

For 10k x 300 matrices:

  • Downsample: Take a random subset of rows (e.g., 10-20% of data) for visualization—this reduces computation time without losing overall trends.
  • Feature Subsetting: Focus on key columns (e.g., top 50 correlated features, or columns relevant to your transformation) instead of plotting all 300.
  • Dimensionality Reduction: Use PCA, t-SNE, or UMAP to condense high-dimensional data into 2D/3D plots for easier comparison.

内容的提问来源于stack exchange,提问作者Omar Ayman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 09:05:40