基于DBSCAN的轨迹路径识别报错求助:TypeError问题排查
Let's walk through exactly why this error is popping up and fix your DBSCAN clustering code for trajectory data.
The Root Cause
That error is telling you that np.radians() is being handed individual float values instead of a structured array it can iterate over. Let's look at how you're prepping your coordinates:
coords = a['Long'] coords['Lat'] = a['Lat'] coords = coords.to_numpy(coords)
Here, you start with a 1D Series of longitude values, then try to tack on latitude values as a new "key" in that Series—this doesn't create the 2D coordinate pairs DBSCAN needs. Instead, you end up with a messy 1D structure, and when np.radians() tries to process it, it's looking for a .radians() method on each float (which doesn't exist).
The Fixed Code
Here's the corrected version of your function, with explanations for each change:
import numpy as np from sklearn.cluster import DBSCAN def clustersDBSCAN(data): # Make a copy of the input to avoid modifying the original DataFrame trajectory_data = data.copy() # Create a proper 2D array of [Lat, Long] pairs (critical for haversine metric!) # Note: Haversine expects coordinates in (latitude, longitude) order coords = trajectory_data[['Lat', 'Long']].to_numpy() kms_per_radian = 6371.0088 # Convert 0.02 km to radians for the epsilon threshold epsilon = 0.02 / kms_per_radian # Fit DBSCAN with radians-converted coordinates db = DBSCAN( eps=epsilon, min_samples=1, algorithm='ball_tree', metric='haversine' ).fit(np.radians(coords)) # Attach cluster labels to the DataFrame trajectory_data['clusters'] = db.labels_ return trajectory_data
Key Fixes Explained
- Proper 2D Coordinate Array: By using
trajectory_data[['Lat', 'Long']].to_numpy(), we create a 2D array where each row is a valid (latitude, longitude) pair. This is exactly whatnp.radians()and the haversine metric expect. - Order Matters: The haversine metric in scikit-learn requires coordinates in
(lat, lon)order—swapping these would lead to incorrect distance calculations, even if the error goes away. - Avoid Modifying Original Data: Using
data.copy()prevents accidental changes to your input DataFrame, which is a safe practice for data processing functions. - Correct
to_numpy()Usage: The original code passedcoordsas an argument toto_numpy(), which is unnecessary—to_numpy()doesn't need that parameter.
Quick Pre-Run Check
Make sure your Lat and Long columns in the input DataFrame are numeric (float or int) types. If they're stored as strings, convert them first with:
import pandas as pd data['Lat'] = pd.to_numeric(data['Lat'], errors='coerce') data['Long'] = pd.to_numeric(data['Long'], errors='coerce')
内容的提问来源于stack exchange,提问作者nicolax9777

