如何在部分缺失的规则网格中插值/外推?Python实现问询
Got it, let's solve this problem step by step. You've got a regular grid stored in a Pandas DataFrame with missing values, and you need to fill those gaps with linear interpolation, plus use nearest-neighbor extrapolation for points outside the grid bounds. Here are two practical approaches to get this done:
Approach 1: Scipy + Pandas (Flexible for Custom Workflows)
First, let's formalize your sample data so we can test things out:
import numpy as np import pandas as pd from scipy.interpolate import griddata, NearestNDInterpolator # Your original grid coordinates x_coords = [0, 1, 2, 3, 4] y_coords = [0.5, 1.5, 2.5, 3.5, 4.5, 5.5] # Full z array (5 rows for x, 6 columns for y) with NaNs z_vals = np.array([ [np.nan, np.nan, 1.5, 2.0, 5.5, 3.5], [np.nan, 1.0, 4.0, 2.5, 4.5, 3.0], [2.0, 0.5, 6.0, 1.5, 3.5, np.nan], [np.nan, 1.5, 4.0, 2.0, np.nan, np.nan], [np.nan, np.nan, 2.0, np.nan, np.nan, 1.0] ]) # Convert to DataFrame (x as index, y as columns) df = pd.DataFrame(z_vals, index=x_coords, columns=y_coords)
Step 1: Extract Valid Points & Build Interpolators
We'll first pull out all the non-NaN points from the grid, then create two interpolators: one for linear interpolation (to fill internal gaps) and one for nearest-neighbor (to handle gaps linear interpolation can't fill, plus extrapolation):
# Generate full grid coordinates (matches the shape of z_vals) X, Y = np.meshgrid(x_coords, y_coords, indexing='ij') # 'ij' keeps x as rows, y as columns # Filter out valid (non-NaN) points valid_mask = ~np.isnan(z_vals) valid_points = np.column_stack((X[valid_mask], Y[valid_mask])) valid_values = z_vals[valid_mask] # Linear interpolation for internal gaps linear_result = griddata(valid_points, valid_values, (X, Y), method='linear') # Nearest-neighbor interpolation for remaining gaps + extrapolation nearest_interpolator = NearestNDInterpolator(valid_points, valid_values) nearest_result = nearest_interpolator(X, Y) # Combine results: use linear where possible, fall back to nearest-neighbor otherwise final_z = np.where(~np.isnan(linear_result), linear_result, nearest_result) # Convert back to DataFrame interpolated_df = pd.DataFrame(final_z, index=x_coords, columns=y_coords)
Step 2: Extrapolate Outside Grid Bounds
To get values for points outside your original grid (like x=-0.5 or y=6.5), just use the nearest_interpolator we built:
# Example: Get extrapolated value for x=-0.5, y=6.0 extrapolated_val = nearest_interpolator(-0.5, 6.0) print(f"Extrapolated value: {extrapolated_val}")
Approach 2: Xarray (Simpler for Regular Grids)
If you're open to using Xarray (a great tool for labeled grid data), this process becomes much more concise. Xarray handles the grid alignment automatically:
import xarray as xr # Convert your DataFrame to an Xarray Dataset ds = xr.Dataset( {'z': (['x', 'y'], z_vals)}, coords={'x': x_coords, 'y': y_coords} ) # First, fill internal gaps with linear interpolation (no extrapolation yet) linear_interp_ds = ds.interpolate_na(dims=['x', 'y'], method='linear') # Then, fill remaining gaps (and extrapolate outside bounds) with nearest-neighbor final_ds = linear_interp_ds.interpolate_na( dims=['x', 'y'], method='nearest', fill_value='extrapolate' ) # Convert back to Pandas DataFrame if needed interpolated_df_xarray = final_ds['z'].to_pandas()
Check the Results
If you print interpolated_df or interpolated_df_xarray, you'll see all the NaN values from your original grid are filled in, and you can safely extrapolate to points outside the original bounds using the nearest-neighbor interpolator.
内容的提问来源于stack exchange,提问作者abenhamou

