如何基于1980-2018年MERRA2 .nc文件绘制CLDHGH变量时间序列图
Got it, let's tackle this step by step. You've already nailed plotting a single day's CLDHGH time series, so scaling this up to the full 1980–2018 period just needs a few key additions: generating the complete date range, looping through each file, aggregating your data, and building a proper time axis for plotting.
First, we need to create all the yyyymmdd date strings that match your filename pattern. In MATLAB, you can do this with datetime utilities:
% Define your start and end dates start_date = datetime(1980, 1, 1); end_date = datetime(2018, 3, 30); % Generate every day in the range all_dates = start_date:days(1):end_date; % Convert dates to the yyyymmdd string format used in your filenames date_strings = datestr(all_dates, 'yyyymmdd');
Next, we’ll iterate over each date, load the corresponding .nc file, apply your 0-value cleanup, and store the data. Note: If you’re working with the full dataset (over 330k hourly timesteps), memory might be an issue—so I’ll include two approaches:
Approach A: Full hourly data (if memory allows)
This stores every hour’s 65x80 grid data for the entire period:
% Base path and file extension for your .nc files file_base = 'D:\cloudnc\MERRA2_100.tavg1_2d_rad_Nx.'; file_ext = '.SUB.nc'; % Initialize an empty array to hold all CLDHGH data cldhigh_full = []; for i = 1:length(date_strings) % Build the full filename current_file = [file_base, date_strings(i), file_ext]; % Skip missing files to avoid errors if ~exist(current_file, 'file') warning('File %s not found—skipping this date.', current_file); continue; end % Read the CLDHGH variable cldhigh = ncread(current_file, 'CLDHGH'); % Apply your existing 0-value handling (replace ... with your actual logic) cldhigh(cldhigh == 0) = NaN; % Example: replace 0 with NaN % Append the day's data to the full dataset cldhigh_full = cat(3, cldhigh_full, cldhigh); end
Approach B: Memory-efficient (compute hourly means on the fly)
If storing the full 3D array is too much, calculate regional hourly means as you go to save memory:
hourly_mean = []; for i = 1:length(date_strings) current_file = [file_base, date_strings(i), file_ext]; if ~exist(current_file, 'file') warning('File %s not found—adding NaNs for missing hours.', current_file); hourly_mean = [hourly_mean, NaN(1,24)]; continue; end cldhigh = ncread(current_file, 'CLDHGH'); cldhigh(cldhigh == 0) = NaN; % Calculate mean across the 65x80 grid for each hour daily_hourly_means = squeeze(mean(mean(cldhigh, 1), 2)); hourly_mean = [hourly_mean, daily_hourly_means]; end
To plot your data with meaningful dates, generate a time vector that matches every hour in your dataset:
full_time_axis = start_date:hours(1):end_date + hours(23);
Choose the plot that fits your needs:
Option 1: Regional hourly mean time series
If you used Approach B:
figure; plot(full_time_axis, hourly_mean, 'LineWidth', 0.5); xlabel('Time'); ylabel('Mean CLDHGH'); title('Hourly Mean CLDHGH (1980-01-01 to 2018-03-30)'); datetick('x', 'yyyy-mm', 'keepticks'); % Format x-axis dates for readability grid on;
Option 2: Single grid point time series
If you used Approach A and want to plot a specific grid cell (e.g., row 20, column 30):
% Extract data for the target grid point grid_point_data = cldhigh_full(20, 30, :); figure; plot(full_time_axis, grid_point_data); xlabel('Time'); ylabel('CLDHGH at Grid (20, 30)'); title('Hourly CLDHGH at Grid Point (20, 30) (1980-01-01 to 2018-03-30)'); datetick('x', 'yyyy-mm', 'keepticks'); grid on;
- Parallelize for speed: If you have MATLAB’s Parallel Computing Toolbox, replace
forwithparforto process multiple files at once. - Log missing files: Add a line to write missing dates to a text file (e.g.,
fprintf(log_file, '%s\n', date_strings(i));) for later troubleshooting. - Chunk processing: If even the full hourly mean array is too big, process the data in monthly chunks instead of all at once.
内容的提问来源于stack exchange,提问作者Vaishali Jain

