Verilog实现CNN中3×3滤波器图像卷积的滑窗问题咨询
Hey there! Let’s dig into your 3x3 convolution data fetch problem for that 5x5 image in Verilog—this is such a common pain point when building CNN hardware, so I’ve got two solid approaches to share with you.
First, let’s recap why this is tricky: when sliding a 3x3 filter with stride 1 over a 5x5 image, each convolution window needs 9 pixels. For example, the first window uses pixels (0,0) to (2,2), the next uses (0,1) to (2,3), and so on until we reach the last window (2,2) to (4,4) (since 5-3+1=3, we end up with a 3x3 output feature map).
The issue is that memory usually stores images row-by-row, so you can’t just read consecutive addresses to get the full 3x3 window. You need to retain previous rows’ data and align it with the current row’s pixels as the filter slides.
The go-to fix here is using a line buffer—two shift registers (or small RAM blocks) that store the previous two rows of the image. Combined with the current row you’re reading from memory, you can easily assemble the 3x3 window as the filter slides.
Here’s a simplified Verilog example tailored to your 5x5 image:
module line_buffer #( parameter IMG_WIDTH = 5, // 5 columns for your 5x5 image parameter DATA_WIDTH = 8 // Assuming 8-bit pixel values )( input clk, input rst_n, input [DATA_WIDTH-1:0] pixel_in, // Pixel read from memory output reg [DATA_WIDTH-1:0] window [2:0][2:0] // 3x3 convolution window ); // Line buffers to store the previous two rows reg [DATA_WIDTH-1:0] line1 [IMG_WIDTH-1:0]; reg [DATA_WIDTH-1:0] line2 [IMG_WIDTH-1:0]; reg [DATA_WIDTH-1:0] current_line [IMG_WIDTH-1:0]; // Counters to track current row/column in the input image reg [2:0] col_cnt; // 0-4 for 5 columns reg [2:0] row_cnt; // 0-4 for 5 rows always @(posedge clk or negedge rst_n) begin if(!rst_n) begin // Reset all counters and buffers col_cnt <= 0; row_cnt <= 0; foreach(line1[i]) line1[i] <= 0; foreach(line2[i]) line2[i] <= 0; foreach(current_line[i]) current_line[i] <= 0; foreach(window[i,j]) window[i][j] <= 0; end else begin // 1. Store incoming pixel in the current row's column current_line[col_cnt] <= pixel_in; // 2. Shift line buffers when we reach the end of a row if(col_cnt == IMG_WIDTH-1) begin col_cnt <= 0; line2 <= line1; // Shift oldest row out line1 <= current_line; // Move current row to line1 row_cnt <= row_cnt + 1; end else begin col_cnt <= col_cnt + 1; end // 3. Assemble the 3x3 window once we have enough data if(row_cnt >= 2 && col_cnt >= 2) begin // Window rows: line2 (oldest), line1, current_line (newest) window[0][0] = line2[col_cnt-2]; window[0][1] = line2[col_cnt-1]; window[0][2] = line2[col_cnt]; window[1][0] = line1[col_cnt-2]; window[1][1] = line1[col_cnt-1]; window[1][2] = line1[col_cnt]; window[2][0] = current_line[col_cnt-2]; window[2][1] = current_line[col_cnt-1]; window[2][2] = current_line[col_cnt]; end end end endmodule
Key Details:
- We wait until we’ve read the first two full rows (
row_cnt >=2) and are at the third column (col_cnt >=2) before starting to output valid windows. - Each time we finish a row, we shift the line buffers to keep the most recent two rows available.
- The window updates every clock cycle once valid, matching your stride 1 requirement.
If you don’t want to use line buffers (e.g., for lower resource usage), you can calculate the memory address for each of the 9 pixels in the current window directly. This works best if your memory supports concurrent reads of multiple addresses.
Here’s how to generate the addresses for your 5x5 image:
module addr_generator #( parameter IMG_WIDTH = 5, parameter DATA_WIDTH = 8 )( input clk, input rst_n, input start_conv, // Signal to start convolution output reg [7:0] addr [8:0], // 9 addresses for the 3x3 window output reg valid_addr // High when addresses are valid ); // Counters for the output feature map (3x3, so 0-2 for rows/columns) reg [2:0] out_row_cnt; reg [2:0] out_col_cnt; always @(posedge clk or negedge rst_n) begin if(!rst_n) begin out_row_cnt <= 0; out_col_cnt <= 0; valid_addr <= 0; foreach(addr[i]) addr[i] <= 0; end else if(start_conv) begin valid_addr <= 1; // Calculate addresses for the current 3x3 window // Top row of the window addr[0] = out_row_cnt * IMG_WIDTH + out_col_cnt; addr[1] = out_row_cnt * IMG_WIDTH + out_col_cnt + 1; addr[2] = out_row_cnt * IMG_WIDTH + out_col_cnt + 2; // Middle row addr[3] = (out_row_cnt + 1) * IMG_WIDTH + out_col_cnt; addr[4] = (out_row_cnt + 1) * IMG_WIDTH + out_col_cnt + 1; addr[5] = (out_row_cnt + 1) * IMG_WIDTH + out_col_cnt + 2; // Bottom row addr[6] = (out_row_cnt + 2) * IMG_WIDTH + out_col_cnt; addr[7] = (out_row_cnt + 2) * IMG_WIDTH + out_col_cnt + 1; addr[8] = (out_row_cnt + 2) * IMG_WIDTH + out_col_cnt + 2; // Update counters for the next window if(out_col_cnt == 2) begin out_col_cnt <= 0; if(out_row_cnt == 2) begin valid_addr <= 0; // Convolution complete end else begin out_row_cnt <= out_row_cnt + 1; end end else begin out_col_cnt <= out_col_cnt + 1; end end else begin valid_addr <= 0; end end endmodule
Pros & Cons:
- Pros: No extra buffer memory needed.
- Cons: Requires your memory to handle 9 concurrent read requests (which can be bandwidth-heavy). Great for small images like your 5x5 case, but less scalable for larger images.
- Remember that your output feature map will be 3x3 (since
input_size - filter_size + 1 = 5-3+1=3), so your counters should stop once they reach 2 (0-indexed). - If you’re using a block RAM for image storage, make sure it’s configured to support the read bandwidth you need (either sequential reads for the line buffer approach, or parallel reads for the address generation method).
内容的提问来源于stack exchange,提问作者ViniLL

