如何高效遍历DataFrame列对计算Frechet距离并生成结果数据集?
Got it, let's fix this manual grind with a clean, efficient approach. Here's how you can automate computing Fréchet distances for all your column pairs and package the results into a usable dataset:
Step 1: Organize Your Columns into Groups
First, we'll group your columns by their base number (e.g., columns 1 and 1.1 go into group 1, 2 and 2.1 into group 2, etc.). This eliminates the need to hardcode column names.
import pandas as pd import numpy as np import similaritymeasures from itertools import combinations # Assuming your DataFrame `df` is already loaded column_groups = {} # Iterate through columns to build groups for col in df.columns: # Extract the base integer ID from column names (handles both "1" and "1.1") base_id = int(str(col).split('.')[0]) if base_id not in column_groups: column_groups[base_id] = [] column_groups[base_id].append(col) # Sort each group's columns to ensure consistent x/y order (e.g., "1" before "1.1") for group_id in column_groups: column_groups[group_id].sort() # Convert groups into a list for easy combination generation group_list = list(column_groups.items()) # Each entry: (group_id, [x_col, y_col])
Step 2: Generate All Column Pair Combinations
Use itertools.combinations to create all unique unordered pairs of groups (avoids redundant calculations like group 1 vs group 2 and group 2 vs group 1). If you need ordered pairs (including reverse), swap combinations with product from itertools.
# Generate all unique unordered group pairs pair_combinations = combinations(group_list, 2)
Step 3: Batch Compute Fréchet Distances
Loop through each pair, compute the distance, and collect results. For large datasets (like 1600 groups), parallel processing will drastically speed things up—here's both sequential and parallel versions:
Sequential Version (Simple for Smaller Datasets)
distance_results = [] for (id1, cols1), (id2, cols2) in pair_combinations: # Prepare P (from first group) x1, y1 = df[cols1[0]], df[cols1[1]] P_final = list(zip(x1, y1)) # Prepare Q (from second group) x2, y2 = df[cols2[0]], df[cols2[1]] Q_final = list(zip(x2, y2)) # Calculate Fréchet distance dist = similaritymeasures.frechet_dist(P_final, Q_final) # Store result distance_results.append({ 'Group_1': id1, 'Group_2': id2, 'Frechet_Distance': dist }) # Convert results to a DataFrame result_df = pd.DataFrame(distance_results)
Parallel Version (Fast for 1600 Groups)
For 1600 groups, you'll have ~1.28 million pairs to compute. Use joblib to parallelize across CPU cores:
from joblib import Parallel, delayed def compute_frechet_pair(pair): (id1, cols1), (id2, cols2) = pair x1, y1 = df[cols1[0]], df[cols1[1]] P_final = list(zip(x1, y1)) x2, y2 = df[cols2[0]], df[cols2[1]] Q_final = list(zip(x2, y2)) return { 'Group_1': id1, 'Group_2': id2, 'Frechet_Distance': similaritymeasures.frechet_dist(P_final, Q_final) } # Run in parallel (n_jobs=-1 uses all available CPU cores) distance_results = Parallel(n_jobs=-1)(delayed(compute_frechet_pair)(pair) for pair in pair_combinations) result_df = pd.DataFrame(distance_results)
Step 4: Optional - Reshape into a Distance Matrix
If you prefer a matrix format where rows and columns are group IDs and values are distances, use pivot:
result_matrix = result_df.pivot(index='Group_1', columns='Group_2', values='Frechet_Distance')
Key Notes
- Validation: Add a check to ensure every group has exactly two columns (in case of missing data):
for group_id, cols in column_groups.items(): if len(cols) != 2: print(f"Warning: Group {group_id} has {len(cols)} columns instead of 2") - Memory: For 1600 groups, the distance matrix will be ~1600x1600 (2.56 million entries), which is manageable in most cases.
内容的提问来源于stack exchange,提问作者Mamed

