能否在Amazon SageMaker自定义处理作业中实现自包含数据可视化?
Instead of relying on regex-based metric tracking (which is impractical for large data matrices), you can embed visualization logic directly into your processing job and store outputs in S3 for easy access. Here are the most straightforward, native approaches:
1. Generate Visualizations in the Processing Script & Save to S3
SageMaker Processing jobs automatically sync files in the /opt/ml/output/data/ directory to your specified S3 bucket. You can use standard plotting libraries (Matplotlib, Seaborn, Plotly) to generate pre- and post-transformation visuals, then save them to this output directory.
Example Code Snippet
import pandas as pd import seaborn as sns import matplotlib.pyplot as plt import os from sklearn.decomposition import PCA # Load raw input data (from processing job's input channel) input_dir = "/opt/ml/input/data/raw" raw_df = pd.read_csv(os.path.join(input_dir, "raw_dataset.csv")) # --- Pre-processing visualization --- # Generate correlation heatmap (subset columns if needed for large datasets) plt.figure(figsize=(14, 10)) sns.heatmap(raw_df.corr(numeric_only=True), annot=False, cmap="viridis") plt.title("Pre-Processing Feature Correlation") # Save to output directory (auto-synced to S3) output_dir = "/opt/ml/output/data/visualizations" os.makedirs(output_dir, exist_ok=True) plt.savefig(os.path.join(output_dir, "pre_correlation_heatmap.png")) plt.close() # Generate PCA plot for high-dimensional data pca = PCA(n_components=2) raw_pca = pca.fit_transform(raw_df.select_dtypes(include='number')) raw_pca_df = pd.DataFrame({"PC1": raw_pca[:,0], "PC2": raw_pca[:,1]}) plt.figure(figsize=(10, 8)) plt.scatter(raw_pca_df["PC1"], raw_pca_df["PC2"], alpha=0.5) plt.title("Pre-Processing PCA Projection") plt.savefig(os.path.join(output_dir, "pre_pca_plot.png")) plt.close() # --- Perform your data transformation here --- transformed_df = raw_df.apply(your_transformation_logic) # --- Post-processing visualization --- plt.figure(figsize=(14, 10)) sns.heatmap(transformed_df.corr(numeric_only=True), annot=False, cmap="viridis") plt.title("Post-Processing Feature Correlation") plt.savefig(os.path.join(output_dir, "post_correlation_heatmap.png")) plt.close() # Post-transformation PCA plot transformed_pca = pca.fit_transform(transformed_df.select_dtypes(include='number')) transformed_pca_df = pd.DataFrame({"PC1": transformed_pca[:,0], "PC2": transformed_pca[:,1]}) plt.figure(figsize=(10, 8)) plt.scatter(transformed_pca_df["PC1"], transformed_pca_df["PC2"], alpha=0.5) plt.title("Post-Processing PCA Projection") plt.savefig(os.path.join(output_dir, "post_pca_plot.png")) plt.close()
When defining your processing job, specify the output S3 path in the ProcessingOutput configuration—all files in /opt/ml/output/data/visualizations will be uploaded there once the job completes.
2. View Visualizations in SageMaker Studio
Once the visuals are in S3, you can access them directly in SageMaker Studio:
- Navigate to your S3 bucket via the Studio file browser, right-click on an image file, and select "Open" to view it.
- For interactive plots (e.g., Plotly HTML files), open them in Studio's built-in web viewer.
Alternatively, use a notebook cell to display images inline:
from IPython.display import Image import sagemaker s3_session = sagemaker.Session().s3_resource bucket = "your-bucket-name" prefix = "processing-job-output/visualizations" # Display pre-processing heatmap pre_heatmap = s3_session.Object(bucket, f"{prefix}/pre_correlation_heatmap.png").get()["Body"].read() Image(pre_heatmap) # Display post-processing PCA plot post_pca = s3_session.Object(bucket, f"{prefix}/post_pca_plot.png").get()["Body"].read() Image(post_pca)
3. Optimizations for Large Datasets
For 10k x 300 matrices:
- Downsample: Take a random subset of rows (e.g., 10-20% of data) for visualization—this reduces computation time without losing overall trends.
- Feature Subsetting: Focus on key columns (e.g., top 50 correlated features, or columns relevant to your transformation) instead of plotting all 300.
- Dimensionality Reduction: Use PCA, t-SNE, or UMAP to condense high-dimensional data into 2D/3D plots for easier comparison.
内容的提问来源于stack exchange,提问作者Omar Ayman

