Python下含间隙异常值的实测与模拟光变曲线对比及拟合问询
Hey there! Let's tackle your two questions about comparing and fitting light curves, tailored to your specific scenario:
Since your observed data has gaps, outliers, and irregular time steps (while the simulation uses fixed steps), you need methods that can handle these mismatches. Here are the most suitable options:
Dynamic Time Warping (DTW)
This is perfect for comparing time series with different lengths and sampling rates. DTW aligns the two curves by "warping" the time axis, effectively ignoring gaps and accommodating irregular steps. You can preprocess the observed data to remove extreme outliers first, then use Python libraries likedtw-pythonto compute the similarity score.Phase-folded statistical comparisons
Light curves are periodic, so first fold both the observed data (using either the simulation's period or an estimated period from your data) and the simulated curve onto a common phase axis. Then:- Calculate RMSE (Root Mean Squared Error) or MAE (Mean Absolute Error) between the folded observed points and the simulated curve (make sure to calibrate magnitudes to the same scale first).
- Use the Spearman correlation coefficient (more robust to outliers than Pearson) to measure the monotonic relationship between the two datasets.
Kolmogorov-Smirnov (K-S) Test
After phase folding, this non-parametric test compares the overall magnitude distributions of the simulated and observed data. It tells you if the two datasets likely come from the same underlying distribution, which is a good measure of overall similarity.
Your brute-force fit idea is valid, but let's weigh its pros and cons, plus explore better alternatives for your 4-parameter model:
Brute-Force Fit: What to know
- Pros: Super straightforward to implement, no need for advanced optimization knowledge. Works well if your parameter ranges are narrow and you can use coarse steps.
- Cons: Gets computationally expensive fast if parameter ranges are wide or you need fine steps—4 parameters with 10 steps each already means 10,000 simulations.
Better alternatives for black-box model fitting
Since your model is generated by external software (a "black box" from Python's perspective), these methods are more efficient:
Latin Hypercube Sampling (LHS) + Brute-Force Hybrid
Instead of grid-searching every possible parameter combination, use LHS to sample parameter points evenly across the entire parameter space. This reduces redundant simulations while ensuring good coverage, making your brute-force approach way more efficient. Python'sscipy.stats.qmchas tools for LHS.Bayesian Optimization
This is ideal for black-box optimization. It uses a probabilistic model (like a Gaussian process) to "learn" which parameter combinations are most likely to minimize your similarity metric (e.g., RMSE). It intelligently picks the next set of parameters to test, drastically cutting down on the number of simulations needed. Try libraries likescikit-optimizeorbayesian-optimizationin Python, paired withsubprocessto automate calling your external software, passing parameters, and reading output curves.Genetic Algorithms
These mimic natural selection to iteratively refine parameter sets. They're great for escaping local minima and work well with black-box models. Thedeaplibrary in Python has tools to implement genetic algorithms for your parameter search.
Quick automation tip
Use Python's subprocess module to script the entire workflow: pass parameters to your external software, run the simulation, export the light curve data, then compute the similarity metric—no more manual parameter input!
From your three plots (simulated curve, observed data, overlay), it's clear that tweaking parameters can align the phase, period, magnitude scale, and even the flare cycles to get a good match. With the right fitting method, you'll get there efficiently.
内容的提问来源于stack exchange,提问作者Wiesel

