如何在DataFrame中按不同步长遍历XY列,计算各2D盒子的速度均值?
Got it, let's expand your 1D logic to 2D box-based averaging smoothly. The core idea is to add a nested loop for the y-axis intervals, and filter your DataFrame on both x and y ranges for each 2D box. Here's how to do it, keeping your existing unit conversion and step sizes intact:
Step 1: Adjust Your Loop Structure
Replace your single while loop with a nested loop (one for x intervals, one for y intervals). We'll also track each box's position so you can map the averages back to their 2D regions later.
import numpy as np import pandas as pd # Your existing variables vels = ... # Your DataFrame with x, y, vx, vy columns steps = 18.18 # X-axis box width (px) steps1 = 36.36 # Y-axis box height (px) conversion_factor = 2.75 # µm/s per px unit # Initialize result lists MeanVx = [] MeanVy = [] MeanVm = [] box_x_centers = [] box_y_centers = [] # Get max bounds for x and y axes max_x = np.ceil(vels["x"].max()) max_y = np.ceil(vels["y"].max()) # Nested loops to iterate over all 2D boxes i = 0.0 while np.round(i) <= max_x: j = 0.0 while np.round(j) <= max_y: # Filter rows that fall within the current 2D box filter_vels = vels[ (vels["x"] >= i) & (vels["x"] <= i + steps) & (vels["y"] >= j) & (vels["y"] <= j + steps1) ] # Calculate averages (handle empty boxes to avoid NaN errors) if not filter_vels.empty: mean_vx = filter_vels["vx"].mean() * conversion_factor mean_vy = filter_vels["vy"].mean() * conversion_factor mean_vm = np.sqrt(mean_vx**2 + mean_vy**2) else: # Fill empty boxes with NaN (or 0 if you prefer) mean_vx = np.nan mean_vy = np.nan mean_vm = np.nan # Append results and box positions MeanVx.append(mean_vx) MeanVy.append(mean_vy) MeanVm.append(mean_vm) box_x_centers.append(i + steps/2) box_y_centers.append(j + steps1/2) j += steps1 i += steps # Optional: Convert results to a DataFrame for easier analysis/visualization results_df = pd.DataFrame({ "x_center": box_x_centers, "y_center": box_y_centers, "mean_vx": MeanVx, "mean_vy": MeanVy, "mean_vm": MeanVm })
Step 2: Faster Alternative with pandas.cut + groupby
If you're working with large datasets, nested loops can be slow. A more efficient approach uses pandas' built-in binning and grouping tools, which are vectorized and run much faster:
# Define bins for x and y axes x_bins = np.arange(0, max_x + steps, steps) y_bins = np.arange(0, max_y + steps1, steps1) # Create labels using box center points (easier to map later) x_labels = x_bins[:-1] + steps/2 y_labels = y_bins[:-1] + steps1/2 # Assign each row to its corresponding x and y bin vels["x_bin"] = pd.cut(vels["x"], bins=x_bins, labels=x_labels, include_lowest=True) vels["y_bin"] = pd.cut(vels["y"], bins=y_bins, labels=y_labels, include_lowest=True) # Group by bins and calculate average velocities grouped = vels.groupby(["x_bin", "y_bin"]).agg( mean_vx=("vx", lambda x: x.mean() * conversion_factor), mean_vy=("vy", lambda x: x.mean() * conversion_factor) ).reset_index() # Calculate mean velocity magnitude grouped["mean_vm"] = np.sqrt(grouped["mean_vx"]**2 + grouped["mean_vy"]**2)
This method automatically handles empty bins (omits them by default; add dropna=False to groupby if you want to keep them as NaN) and returns a clean DataFrame with all 2D average values tied to their box positions.
Key Notes
- Empty Boxes: If some 2D boxes have no data points, the loop version uses
np.nan(you can swap this for 0 if needed). Thegroupbyversion skips empty bins by default. - Visualization: With the
results_dforgroupedDataFrame, you can easily create 2D heatmaps of mean velocity using libraries likeseaborn.heatmap(reshape the data withpivotfirst to get a matrix format).
内容的提问来源于stack exchange,提问作者Edouardo

