基于CSV文件的正态分布PDF:R与Python实现对比及优化问询
Hey there! Let's break down how to cleanly replicate your R workflow in Python—ditching clunky PDF library workarounds for streamlined, idiomatic code that's easier to maintain and more accurate. Here's a step-by-step guide tailored to your needs:
First, we'll use pandas to load your CSV and keep only columns 2 and 3 (remember Python uses 0-indexing, so we target positions 1 and 2). This mirrors R's read.csv()[,2:3] behavior but with cleaner syntax:
import pandas as pd # Load your CSV and retain only columns 2 & 3 (adjust filename as needed) df = pd.read_csv("your_data.csv").iloc[:, 1:3] # Optional: Rename columns for clarity (match your actual column names if needed) df.columns = ["variable_a", "variable_b"]
Instead of manually computing PDFs with generic libraries, use scipy.stats—it's purpose-built for statistical distributions, so you avoid calculation errors and save time. We'll generate a smooth range of x-values first, then compute PDFs for both columns:
from scipy.stats import norm, t import numpy as np # Generate a dense range of x-values spanning your data's min/max x = np.linspace(df.min().min(), df.max().max(), 1000) # Normal distribution PDFs (using mean/std from each column) norm_pdf_a = norm.pdf(x, loc=df["variable_a"].mean(), scale=df["variable_a"].std()) norm_pdf_b = norm.pdf(x, loc=df["variable_b"].mean(), scale=df["variable_b"].std()) # T-distribution PDFs (degrees of freedom = n-1, standard R convention) t_df_a = len(df["variable_a"]) - 1 t_pdf_a = t.pdf(x, df=t_df_a, loc=df["variable_a"].mean(), scale=df["variable_a"].std()) t_df_b = len(df["variable_b"]) - 1 t_pdf_b = t.pdf(x, df=t_df_b, loc=df["variable_b"].mean(), scale=df["variable_b"].std())
Use matplotlib and seaborn to create clean, publication-ready plots—similar to R's ggplot or base graphics, but with flexible customization. We'll make side-by-side plots for normal and t-distribution comparisons:
import matplotlib.pyplot as plt import seaborn as sns # Set a clean, professional plot style sns.set_style("whitegrid") # Create a figure with 2 side-by-side subplots fig, (ax_norm, ax_t) = plt.subplots(1, 2, figsize=(14, 6)) # Plot Normal PDFs (optional: add histograms of raw data for context) ax_norm.plot(x, norm_pdf_a, label=f"Normal - {df.columns[0]}", color="#1f77b4", linewidth=2) ax_norm.plot(x, norm_pdf_b, label=f"Normal - {df.columns[1]}", color="#ff7f0e", linewidth=2) # Add histograms (density=True to match PDF scale) ax_norm.hist(df["variable_a"], density=True, alpha=0.3, color="#1f77b4") ax_norm.hist(df["variable_b"], density=True, alpha=0.3, color="#ff7f0e") ax_norm.set_title("Normal Distribution PDF Comparison", fontsize=12) ax_norm.set_xlabel("Value", fontsize=10) ax_norm.set_ylabel("Probability Density", fontsize=10) ax_norm.legend() # Plot T-Distributions ax_t.plot(x, t_pdf_a, label=f"T-dist (df={t_df_a}) - {df.columns[0]}", color="#1f77b4", linewidth=2, linestyle="--") ax_t.plot(x, t_pdf_b, label=f"T-dist (df={t_df_b}) - {df.columns[1]}", color="#ff7f0e", linewidth=2, linestyle="--") ax_t.set_title("T-Distribution PDF Comparison", fontsize=12) ax_t.set_xlabel("Value", fontsize=10) ax_t.set_ylabel("Probability Density", fontsize=10) ax_t.legend() # Adjust layout to prevent label overlap plt.tight_layout() plt.show()
Why This Is Better Than Your Previous Approach:
- Statistical Accuracy:
scipy.statsuses industry-standard implementations of distributions, so no manual math errors. - Code Maintainability: Pandas and scipy are standard tools in Python data science—other developers will immediately understand your code.
- Flexibility: You can easily tweak plots (colors, styles, annotations) or extend the analysis (e.g., add confidence intervals) without rewriting core logic.
内容的提问来源于stack exchange,提问作者Moritz

