如何查找Pandas DataFrame中与其他行总距离最小的行索引
Unnamed: 0 Value for Rows with Minimum Total Distance Hey there! Let's walk through how to identify which row(s) have the smallest total distance to all other rows, then map that back to the Unnamed: 0 column in your DataFrame.
Step 1: Calculate Total Distance for Each Row
First, we'll sum up the distances in each row of your distance_matrix—this gives us the total distance each row has to every other row. We can use NumPy's sum function for this:
import numpy as np row_total_distances = np.sum(distance_matrix, axis=1)
For your example, this will output array([6., 3., 3.])—the first row has a total distance of 6, while rows 1 and 2 have a total of 3 each.
Step 2: Locate Rows with the Minimum Total Distance
Next, we'll find the indices of all rows that have the smallest total distance. We can use np.where to get all matching indices (in case multiple rows tie for the minimum):
min_dist_indices = np.where(row_total_distances == row_total_distances.min())[0]
This returns array([1, 2]) for your data, meaning rows 1 and 2 (0-indexed) are the ones with the smallest total distance.
Step 3: Get Corresponding Unnamed: 0 Values
Finally, we'll pull the Unnamed: 0 values from these rows in your original DataFrame. Using unique() ensures we don't get duplicate values if multiple rows share the same index:
result = df.loc[min_dist_indices, 'Unnamed: 0'].unique()
In your case, this outputs array([-527.])—since both rows 1 and 2 have -527.0 in Unnamed: 0.
Simplified Version for Single Minimum Row
If you only need the first row that has the minimum total distance, you can use np.argmin instead:
min_row_idx = np.argmin(row_total_distances) result = df.loc[min_row_idx, 'Unnamed: 0']
This will return -527.0 directly.
Full Working Code
Here's the complete code put together, using your sample data:
import numpy as np import pandas as pd import scipy.spatial.distance as ssd import mlpy # Your original DataFrame df = pd.DataFrame({ 'Unnamed: 0': [-527.0, -527.0, -527.0], 0: [143.0, 143.0, 143.0], 1: [143.0, 143.0, 143.0], 2: [248.0, 248.0, 248.0], 3: [-952.0, -955.0, -955.0] }) # Calculate distance matrix (your existing code) dm = ssd.pdist(df, lambda u,v: mlpy.dtw_std(pd.Series(u).dropna().values.tolist(), pd.Series(v).dropna().values.tolist(), dist_only=True)) distance_matrix = ssd.squareform(dm) # Find the Unnamed: 0 value(s) for rows with minimum total distance row_total_distances = np.sum(distance_matrix, axis=1) min_dist_indices = np.where(row_total_distances == row_total_distances.min())[0] result = df.loc[min_dist_indices, 'Unnamed: 0'].unique() print(result) # Output: [-527.]
内容的提问来源于stack exchange,提问作者user3806649

