Python处理GPS数据:异常值去除与曲线拟合故障排查
Hey there! Let's walk through each of your GPS path processing issues one by one and get them sorted out:
1. Why is the spline fitting through every point (including outliers)?
The Ramer-Douglas-Peucker (RDP) algorithm is for line simplification, not outlier detection. It preserves the overall shape of your path but doesn't filter out random, off-path points. To get rid of outliers, you need an extra step to identify and remove them first:
Quick Outlier Filtering Method
Calculate the distance from each original GPS point to your RDP-simplified line, then drop points that are too far away:
from shapely.geometry import Point, LineString # Convert RDP-simplified points to a line object rdp_line = LineString(simple_coords) # Compute distance from each original point to the RDP line coords['distance_to_rdp'] = coords.apply( lambda row: Point(row['lon'], row['lat']).distance(rdp_line), axis=1 ) # Set a threshold (adjust based on your data's scale—1e-4 works for most GPS coords) outlier_threshold = 1e-4 clean_coords = coords[coords['distance_to_rdp'] < outlier_threshold] # Re-run RDP on the cleaned dataset simple_clean_coords = rdp(clean_coords[['lon', 'lat']].values, epsilon=1e-4)
Now your simplified dataset will exclude those off-path outliers, so your spline won't be forced to fit them.
2. Why does adjusting s make the fit line disappear?
Two parameter mistakes are causing this:
t=10: Manually specifying the number of spline nodes disrupts the algorithm's automatic node selection logic—remove this parameter entirely.sscale mismatch: GPS coordinates are small (decimal degrees with 4-6 decimal places), sos(the maximum allowed sum of squared errors) needs to be tiny. If you setstoo large (like 1), the algorithm will over-smooth your path to the point of erasing it.
Fixed Spline Fitting Code
# Use the cleaned, simplified points for fitting dat_clean = np.array([(x, y) for x, y in zip(simple_clean_coords[:, 0], simple_clean_coords[:, 1])]) # Adjust s to a small value (start with 1e-6) and remove t parameter tck, u = splprep(dat_clean.T, u=None, s=1e-6, per=1) u_new = np.linspace(u.min(), u.max(), 1000) x_new, y_new = splev(u_new, tck, der=0) # Plot results plt.plot(dat_clean[:, 0], dat_clean[:, 1], 'ro', label='Clean Simplified Points') plt.plot(x_new, y_new, 'b--', label='Smoothed B-Spline') plt.legend() plt.show()
Tweak s between 1e-7 and 1e-5 to find the perfect balance between smoothness and path accuracy.
3. Do I have to use Mean Squared Error (MSE) for error calculation?
Absolutely not! MSE is just one option—choose the metric that fits your use case:
- Mean Absolute Error (MAE): More robust to remaining outliers than MSE, and easier to interpret.
- Maximum Error: Tells you the worst-case deviation between your original and fitted path, great for strict accuracy requirements.
- Real-World Distance Error: Since GPS uses spherical coordinates, Euclidean distance doesn't reflect actual ground distance. Use the Haversine formula to calculate errors in meters for real-world relevance.
Example: Calculate Errors in Meters
def haversine(lon1, lat1, lon2, lat2): # Convert degrees to radians lon1, lat1, lon2, lat2 = map(np.radians, [lon1, lat1, lon2, lat2]) dlon = lon2 - lon1 dlat = lat2 - lat1 a = np.sin(dlat/2)**2 + np.cos(lat1) * np.cos(lat2) * np.sin(dlon/2)**2 c = 2 * np.arcsin(np.sqrt(a)) earth_radius_m = 6371000 # Earth's radius in meters return c * earth_radius_m # Create a LineString for the fitted spline fit_line = LineString(np.column_stack((x_new, y_new))) # Calculate actual distance from each cleaned point to the fitted line clean_coords['fit_error_m'] = clean_coords.apply( lambda row: haversine( row['lon'], row['lat'], *fit_line.interpolate(fit_line.project(Point(row['lon'], row['lat']))).coords[0] ), axis=1 ) # Compute different error metrics mae_m = clean_coords['fit_error_m'].mean() max_error_m = clean_coords['fit_error_m'].max() print(f"Average Error: {mae_m:.2f} meters | Maximum Error: {max_error_m:.2f} meters")
内容的提问来源于stack exchange,提问作者surajr

