You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python处理GPS数据:异常值去除与曲线拟合故障排查

Troubleshooting Your GPS Path Processing Workflow

Hey there! Let's walk through each of your GPS path processing issues one by one and get them sorted out:

1. Why is the spline fitting through every point (including outliers)?

The Ramer-Douglas-Peucker (RDP) algorithm is for line simplification, not outlier detection. It preserves the overall shape of your path but doesn't filter out random, off-path points. To get rid of outliers, you need an extra step to identify and remove them first:

Quick Outlier Filtering Method

Calculate the distance from each original GPS point to your RDP-simplified line, then drop points that are too far away:

from shapely.geometry import Point, LineString

# Convert RDP-simplified points to a line object
rdp_line = LineString(simple_coords)

# Compute distance from each original point to the RDP line
coords['distance_to_rdp'] = coords.apply(
    lambda row: Point(row['lon'], row['lat']).distance(rdp_line),
    axis=1
)

# Set a threshold (adjust based on your data's scale—1e-4 works for most GPS coords)
outlier_threshold = 1e-4
clean_coords = coords[coords['distance_to_rdp'] < outlier_threshold]

# Re-run RDP on the cleaned dataset
simple_clean_coords = rdp(clean_coords[['lon', 'lat']].values, epsilon=1e-4)

Now your simplified dataset will exclude those off-path outliers, so your spline won't be forced to fit them.

2. Why does adjusting s make the fit line disappear?

Two parameter mistakes are causing this:

  • t=10: Manually specifying the number of spline nodes disrupts the algorithm's automatic node selection logic—remove this parameter entirely.
  • s scale mismatch: GPS coordinates are small (decimal degrees with 4-6 decimal places), so s (the maximum allowed sum of squared errors) needs to be tiny. If you set s too large (like 1), the algorithm will over-smooth your path to the point of erasing it.

Fixed Spline Fitting Code

# Use the cleaned, simplified points for fitting
dat_clean = np.array([(x, y) for x, y in zip(simple_clean_coords[:, 0], simple_clean_coords[:, 1])])

# Adjust s to a small value (start with 1e-6) and remove t parameter
tck, u = splprep(dat_clean.T, u=None, s=1e-6, per=1)
u_new = np.linspace(u.min(), u.max(), 1000)
x_new, y_new = splev(u_new, tck, der=0)

# Plot results
plt.plot(dat_clean[:, 0], dat_clean[:, 1], 'ro', label='Clean Simplified Points')
plt.plot(x_new, y_new, 'b--', label='Smoothed B-Spline')
plt.legend()
plt.show()

Tweak s between 1e-7 and 1e-5 to find the perfect balance between smoothness and path accuracy.

3. Do I have to use Mean Squared Error (MSE) for error calculation?

Absolutely not! MSE is just one option—choose the metric that fits your use case:

  • Mean Absolute Error (MAE): More robust to remaining outliers than MSE, and easier to interpret.
  • Maximum Error: Tells you the worst-case deviation between your original and fitted path, great for strict accuracy requirements.
  • Real-World Distance Error: Since GPS uses spherical coordinates, Euclidean distance doesn't reflect actual ground distance. Use the Haversine formula to calculate errors in meters for real-world relevance.

Example: Calculate Errors in Meters

def haversine(lon1, lat1, lon2, lat2):
    # Convert degrees to radians
    lon1, lat1, lon2, lat2 = map(np.radians, [lon1, lat1, lon2, lat2])
    dlon = lon2 - lon1
    dlat = lat2 - lat1
    a = np.sin(dlat/2)**2 + np.cos(lat1) * np.cos(lat2) * np.sin(dlon/2)**2
    c = 2 * np.arcsin(np.sqrt(a))
    earth_radius_m = 6371000  # Earth's radius in meters
    return c * earth_radius_m

# Create a LineString for the fitted spline
fit_line = LineString(np.column_stack((x_new, y_new)))

# Calculate actual distance from each cleaned point to the fitted line
clean_coords['fit_error_m'] = clean_coords.apply(
    lambda row: haversine(
        row['lon'], row['lat'],
        *fit_line.interpolate(fit_line.project(Point(row['lon'], row['lat']))).coords[0]
    ),
    axis=1
)

# Compute different error metrics
mae_m = clean_coords['fit_error_m'].mean()
max_error_m = clean_coords['fit_error_m'].max()
print(f"Average Error: {mae_m:.2f} meters | Maximum Error: {max_error_m:.2f} meters")

内容的提问来源于stack exchange,提问作者surajr

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:00:56