Python&Scipy:如何用分箱角度数据估计Von Mises尺度因子?
Hey there! Let's work through fixing that scaling mismatch between your Von Mises PDF and binned angle data. First, a quick clarification: the Von Mises distribution doesn't natively have a "scale parameter" like some linear distributions—what you're seeing is a mismatch between the PDF's probability density (which integrates to 1) and your binned data's count-based histogram height. Here are two solid approaches to fix this:
Approach 1: Scale the PDF Directly (Best for Large Weights)
Since you're already using circstats.vmpar() to get loc (μ) and kappa (κ) with weighted data, you just need to calculate a scaling factor to align the PDF with your binned histogram. Here's how:
Step-by-Step Code
import numpy as np import circstats as cs import matplotlib.pyplot as plt # Your binned data (replace with your actual values) binned_angles = np.array([0, np.pi/4, np.pi/2, 3*np.pi/4, np.pi]) bin_weights = np.array([5, 10, 15, 10, 5]) # Estimate Von Mises parameters using weighted binned data loc, kappa = cs.vmpar(binned_angles, weights=bin_weights) # Generate a smooth angle sequence for plotting the PDF x = np.linspace(-np.pi, np.pi, 1000) # Calculate raw Von Mises PDF (probability density, integrates to 1) raw_pdf = cs.vmpdf(x, loc=loc, kappa=kappa) # Calculate scaling factor to match histogram height: total_samples = np.sum(bin_weights) bin_width = binned_angles[1] - binned_angles[0] # Assumes uniform bins scale_factor = total_samples * bin_width # Scale the PDF to align with your binned data scaled_pdf = raw_pdf * scale_factor # Plot the results plt.bar(binned_angles, bin_weights, width=bin_width, alpha=0.5, label='Binned Angle Data') plt.plot(x, scaled_pdf, color='orange', label='Scaled Von Mises PDF') plt.legend() plt.title('Binned Data vs. Scaled Von Mises PDF') plt.show()
For Non-Uniform Bins
If your bins aren't equal width, adjust the histogram to use density instead of raw counts, and scale the PDF by total samples to match the histogram's total area:
# Example non-uniform bins bin_edges = np.array([-np.pi/4, 0, np.pi/4, np.pi/2, 3*np.pi/4, np.pi, 5*np.pi/4]) binned_angles = (bin_edges[:-1] + bin_edges[1:])/2 # Calculate bin centers bin_weights = np.array([5, 10, 15, 10, 5, 3]) bin_widths = bin_edges[1:] - bin_edges[:-1] # Estimate parameters loc, kappa = cs.vmpar(binned_angles, weights=bin_weights) # Generate PDF x = np.linspace(-np.pi, np.pi, 1000) raw_pdf = cs.vmpdf(x, loc=loc, kappa=kappa) # Scale factor = total samples (matches PDF area to histogram area) scale_factor = np.sum(bin_weights) scaled_pdf = raw_pdf * scale_factor # Plot density-based histogram and scaled PDF plt.bar(binned_angles, bin_weights / bin_widths, width=bin_widths, alpha=0.5, label='Binned Data (Density)') plt.plot(x, scaled_pdf, color='orange', label='Scaled Von Mises PDF') plt.legend() plt.show()
Approach 2: Expand Binned Data for Scipy's vonmises.fit()
If you prefer using Scipy's fitting function (even though it doesn't support weights directly), you can "unpack" your binned data into a pseudo-raw dataset by repeating each bin's angle by its weight count. This works well if your weights aren't extremely large:
from scipy.stats import vonmises # Expand binned data into individual sample points expanded_angles = np.repeat(binned_angles, bin_weights) # Fit Von Mises parameters (Scipy returns mu and kappa; no native scale parameter) mu, kappa = vonmises.fit(expanded_angles, f0=1) # f0 = initial guess for kappa # Generate and scale the PDF the same way as Approach 1 raw_pdf_scipy = vonmises.pdf(x, kappa, loc=mu) scaled_pdf_scipy = raw_pdf_scipy * total_samples * bin_width # Plot to compare plt.bar(binned_angles, bin_weights, width=bin_width, alpha=0.5, label='Binned Data') plt.plot(x, scaled_pdf_scipy, color='green', label='Scipy-Fitted Scaled PDF') plt.legend() plt.show()
Key Takeaways
- The "scale parameter" you're missing is just a manual scaling factor to bridge the gap between the PDF's probability density and your binned data's count/density values.
circstats.vmpar()is ideal for weighted binned data—stick with it for parameter estimation, then scale the PDF using total samples and bin width.- Unpacking binned data works for Scipy's fit, but avoid it if you have very large weights (it will create an unnecessarily huge array).
内容的提问来源于stack exchange,提问作者nn123

