GeoPandas中基于列值大小关系的多边形差集计算优化方法问询
Hey there! When you're dealing with a large number of unique value entries, looping through each one can really drag down performance. Let's replace that loop with a vectorized approach using GeoPandas' built-in tools to make this operation way faster.
Background Recap
First, let's confirm the setup—we have a GeoDataFrame with polygons and a value column, and we need to subtract all polygons with a higher value from each polygon in the same value group.
Optimized Vectorized Implementation
Here's how to do it without explicit loops:
import pandas as pd import geopandas as gpd from shapely.geometry import Polygon # Initial data (same as your setup) geoms = gpd.GeoSeries([ Polygon([(0, 0), (2, 0), (2, 2), (0, 2)]), Polygon([(1, 1), (3, 1), (3, 3), (1, 3)]), Polygon([(0, 0), (3, 0), (3, 3), (0, 3)]), ]) gdf = gpd.GeoDataFrame(geometry=geoms) gdf["value"] = [3, 2, 1] # Step 1: Sort the GeoDataFrame by value in descending order sorted_gdf = gdf.sort_values("value", ascending=False).reset_index(drop=True) # Step 2: Compute cumulative union of geometries for all higher values # Expanding apply builds the union of all previous (higher value) geometries # Shift(1) ensures we exclude the current row's geometry from the union sorted_gdf["union_above"] = sorted_gdf["geometry"].expanding().apply( lambda x: x.unary_union ).shift(1) # Step 3: Calculate the difference for each geometry # For the highest value, union_above is NaN—we just keep the original geometry sorted_gdf["result_geometry"] = sorted_gdf.apply( lambda row: row["geometry"].difference(row["union_above"]) if pd.notna(row["union_above"]) else row["geometry"], axis=1 ) # Optional: Restore original order if needed result_gdf = sorted_gdf.sort_values("value", ascending=True).reset_index(drop=True) # Plot to verify import matplotlib.pyplot as plt for idx, row in result_gdf.iterrows(): gpd.GeoSeries(row["result_geometry"]).plot(alpha=0.5) plt.xlim(0, 3) plt.ylim(0, 3) plt.title(f"Value = {row['value']}") plt.show()
How This Works
- Sorting: By sorting in descending order of
value, we can use an expanding window to build up the union of all polygons with higher values as we go down the rows. - Cumulative Union: The
expanding().apply()call creates a running union of all geometries we've processed so far (which are the ones with highervalue). Theshift(1)moves this union down one row, so each row gets the union of only geometries with higher values (not its own). - Vectorized Difference: Using
apply()across rows (with a conditional for the highest value) lets us compute the difference in a single pass, no loops required.
Performance Benefit
This approach avoids the overhead of repeated loc calls and loop iterations. For datasets with many unique value entries, this will be significantly faster than the original loop-based method—GeoPandas handles the heavy lifting with optimized spatial operations under the hood.
内容的提问来源于stack exchange,提问作者andhuang

