使用Matplotlib处理大数组meshgrid时触发内存错误的解决办法
Hey there, that memory error makes total sense—let's break down why it's happening and how to fix it.
Your grid size is massive: 12720 x 10117 equals roughly 128 million points. Storing two full float arrays (x and y) for that many points would take around 2GB of RAM (each float is 8 bytes), and flattening/stacking adds even more overhead. No wonder you're hitting a memory limit.
Here are several practical ways to solve this without loading all points into memory at once:
1. Process Points Row-by-Row (Lowest Memory Footprint)
Instead of generating the entire meshgrid upfront, handle each row individually. This way, you only ever hold one row's worth of points in memory at a time.
Code Example:
import numpy as np from matplotlib.path import Path nx, ny = 12720, 10117 sar_ver = ... # Your path vertices here # Initialize an empty boolean array to store results grid = np.zeros((ny, nx), dtype=bool) # Loop through each row (y-coordinate) for y_idx in range(ny): # Generate all x values for this row x_coords = np.arange(nx) # Create points for this row: shape (nx, 2) row_points = np.column_stack((x_coords, np.full(nx, y_idx))) # Check which points are inside the path grid[y_idx] = Path(sar_ver).contains_points(row_points)
This uses just a few hundred KB of extra memory per iteration (since row_points is only 12720x2 elements), which is totally manageable.
2. Chunk Processing (Balance Speed and Memory)
If row-by-row feels too slow, you can process multiple rows at once in chunks. Adjust the chunk size based on how much RAM you have available (e.g., 100-1000 rows per chunk):
Code Example:
import numpy as np from matplotlib.path import Path nx, ny = 12720, 10117 sar_ver = ... # Your path vertices here chunk_size = 200 # Tweak this based on your system's memory grid = np.zeros((ny, nx), dtype=bool) # Iterate over chunks of rows for start_y in range(0, ny, chunk_size): end_y = min(start_y + chunk_size, ny) num_rows = end_y - start_y # Generate x and y coordinates for the chunk x_chunk = np.tile(np.arange(nx), num_rows) y_chunk = np.repeat(np.arange(start_y, end_y), nx) # Create points array for the chunk chunk_points = np.column_stack((x_chunk, y_chunk)) # Check containment and reshape back to chunk dimensions chunk_results = Path(sar_ver).contains_points(chunk_points).reshape(num_rows, nx) # Save results to the grid grid[start_y:end_y] = chunk_results
Larger chunks mean fewer iterations, so this will be faster than row-by-row, but uses more memory per chunk. Find the sweet spot for your system.
3. Use Shapely for Efficient Geometry Operations
Shapely is a dedicated geometry library that's optimized for these kinds of spatial checks. It can handle vectorized operations well and often has better performance than matplotlib's Path for large datasets.
Code Example:
import numpy as np from shapely.geometry import Polygon, MultiPoint nx, ny = 12720, 10117 sar_ver = ... # Your path vertices here # Convert your path to a Shapely Polygon polygon = Polygon(sar_ver) grid = np.zeros((ny, nx), dtype=bool) # Process each row (or use chunks as above) for y_idx in range(ny): x_coords = np.arange(nx) # Create a MultiPoint object for the row row_points = MultiPoint(np.column_stack((x_coords, np.full(nx, y_idx)))) # Check which points are inside the polygon grid[y_idx] = np.array([point.within(polygon) for point in row_points])
Shapely's within method is optimized, so this might be faster than matplotlib's contains_points for large batches.
4. Optimize Your Path/Polygon
If your sar_ver path has a ton of vertices, simplifying it can drastically speed up containment checks. Both matplotlib and Shapely have built-in simplification methods:
Matplotlib Path Simplification:
simplified_path = Path(sar_ver).simplify(tolerance=0.5) # Adjust tolerance to balance accuracy and speed
Shapely Polygon Simplification:
simplified_polygon = polygon.simplify(tolerance=0.5)
The tolerance parameter controls how much the shape is simplified—higher values mean fewer vertices and faster checks, but slightly less accuracy.
Final Notes
- The row-by-row approach is the safest if you're tight on memory.
- Chunk processing is a good middle ground for speed vs memory.
- Shapely is worth considering if you do a lot of spatial operations—it's more feature-rich than matplotlib's Path module.
内容的提问来源于stack exchange,提问作者Atihska

