如何用Python绘制双平台因子对比散点图(多子图/单文件)
Solution for Comparing NP and Arg Platform Metrics with Scatter Plots
Test Data Preparation
First, generate synthetic test data matching your requirements (100 samples, 9 pairs of metrics):
import pandas as pd import numpy as np # Set random seed for reproducibility np.random.seed(42) # Define 9 metric pairs metrics = ['yield', 'mean', 'std', 'max', 'min', 'median', 'q1', 'q3', 'count'] # Create DataFrame with NP and Arg platform metrics data = {f'{metric}_NP': np.random.normal(loc=50, scale=10, size=100) for metric in metrics} data.update({f'{metric}_Arg': np.random.normal(loc=52, scale=9, size=100) for metric in metrics}) df = pd.DataFrame(data)
1. Combined Subplots in Single Figure
This code creates a 3×3 grid of scatter plots to compare all metric pairs in one figure:
import matplotlib.pyplot as plt # Set up figure and subplot grid fig, axes = plt.subplots(nrows=3, ncols=3, figsize=(15, 12)) axes = axes.flatten() # Convert to 1D array for easy iteration # Define distinct colors for each platform platform_colors = {'NP': '#1f77b4', 'Arg': '#ff7f0e'} # Iterate through each metric pair for idx, metric in enumerate(metrics): ax = axes[idx] # Plot NP vs Arg values for each sample ax.scatter(df[f'{metric}_NP'], df[f'{metric}_Arg'], color=platform_colors['NP'], alpha=0.6, label='NP vs Arg') # Add diagonal reference line (perfect correlation) min_val = min(df[f'{metric}_NP'].min(), df[f'{metric}_Arg'].min()) max_val = max(df[f'{metric}_NP'].max(), df[f'{metric}_Arg'].max()) ax.plot([min_val, max_val], [min_val, max_val], 'k--', alpha=0.5) # Set axis ranges with padding ax.set_xlim(min_val - 5, max_val + 5) ax.set_ylim(min_val - 5, max_val + 5) # Configure labels and title ax.set_xlabel(f'{metric}_NP') ax.set_ylabel(f'{metric}_Arg') ax.set_title(f'{metric} Comparison') ax.legend() # Adjust layout to avoid overlapping elements plt.tight_layout() plt.savefig('combined_metric_comparison.png', dpi=100) plt.show()
2. Individual Scatter Plots (Saved Separately)
This code generates and saves each metric comparison as an independent file:
for metric in metrics: plt.figure(figsize=(8, 6)) # Plot data points plt.scatter(df[f'{metric}_NP'], df[f'{metric}_Arg'], color=platform_colors['NP'], alpha=0.6, label='NP vs Arg') # Add reference diagonal line min_val = min(df[f'{metric}_NP'].min(), df[f'{metric}_Arg'].min()) max_val = max(df[f'{metric}_NP'].max(), df[f'{metric}_Arg'].max()) plt.plot([min_val, max_val], [min_val, max_val], 'k--', alpha=0.5) # Set axis ranges plt.xlim(min_val - 5, max_val + 5) plt.ylim(min_val - 5, max_val + 5) # Configure plot labels and title plt.xlabel(f'{metric}_NP') plt.ylabel(f'{metric}_Arg') plt.title(f'{metric}: NP vs Arg Platform Comparison') plt.legend() # Save to file and close figure plt.savefig(f'{metric}_np_vs_arg.png', dpi=100) plt.close()
Key Details
- Axis Range Control: Calculates minimum and maximum values for each metric, then adds padding to ensure all data points are visible.
- Color Differentiation: Uses high-contrast blue and orange to clearly distinguish platform comparisons.
- Reference Line: The dashed diagonal line helps quickly assess how closely each sample's values align between the two platforms.
- Reproducibility: The fixed random seed ensures test data remains consistent across runs.
内容的提问来源于stack exchange,提问作者user3214212
相关产品推荐
相关产品推荐

