如何将多函数C代码改写为两个带输入输出的函数(适配Zynq+Vivado HLS)
Got it, let's walk through exactly how to split your single C file into two functions tailored for Zynq ZC706's CPU-FPGA setup—one that Vivado HLS can synthesize to hardware, and another that runs on the ARM core. The key here is clean input/output boundaries to make intercommunication smooth.
First, let's align on what goes where:
- FPGA-side function: Put all compute-heavy, parallelizable, latency-sensitive logic here. This needs to be Vivado HLS synthesizable (no dynamic memory, recursion, or complex runtime branching unless optimized with HLS pragmas).
- CPU-side function: Handle control flow, peripheral interactions (like reading sensors or sending data over Ethernet), dynamic decision-making, and orchestrating the FPGA accelerator.
- Hard rule: No shared global variables—all data passed between the two functions must go through explicit input/output parameters (this maps directly to Zynq's AXI interface for CPU-FPGA communication).
Let’s say your original C file has two logical blocks: data preprocessing and a compute-intensive filter. Here’s how to split them:
2.1 FPGA-Side Function (HLS-Compatible)
This function will get synthesized to VHDL/Verilog for the FPGA. We’ll use Xilinx’s HLS-specific data types to ensure hardware compatibility:
#include "ap_int.h" #include "ap_fixed.h" // Match data widths to Zynq's AXI bus (32-bit is standard for most setups) #define DATA_WIDTH 32 typedef ap_uint<DATA_WIDTH> raw_data_t; typedef ap_fixed<16, 8> filtered_data_t; // 16-bit fixed-point: 8 integer, 8 fractional bits // FPGA accelerator function: synthesizable with Vivado HLS void fpga_signal_filter( raw_data_t input_samples[1024], // Input: raw data from CPU filtered_data_t output_samples[1024], // Output: filtered data back to CPU int sample_count, // Control: number of samples to process bool enable_pipeline // Control: enable HLS pipeline optimization ) { // Map parameters to Zynq's AXI interfaces #pragma HLS INTERFACE s_axilite port=return bundle=CTRL_BUS #pragma HLS INTERFACE m_axi port=input_samples bundle=DATA_BUS depth=1024 #pragma HLS INTERFACE m_axi port=output_samples bundle=DATA_BUS depth=1024 #pragma HLS INTERFACE s_axilite port=sample_count bundle=CTRL_BUS #pragma HLS INTERFACE s_axilite port=enable_pipeline bundle=CTRL_BUS // Compute-heavy filter logic (example: moving average) for (int i = 0; i < sample_count; i++) { #pragma HLS PIPELINE II=1 if enable_pipeline // Optimize for parallelism filtered_data_t sum = 0; // 5-tap moving average for (int j = max(0, i-2); j <= min(sample_count-1, i+2); j++) { sum += filtered_data_t(input_samples[j]) * 0.2; } output_samples[i] = sum; } }
Key Notes for HLS Compatibility:
- Use
ap_uint/ap_fixedinstead of standard C types—these eliminate ambiguity for hardware synthesis. - The
#pragma HLS INTERFACEdirectives tell Vivado how to map the function to Zynq’s AXI buses (m_axifor high-speed data,s_axilitefor control signals). - Avoid any runtime decisions that can’t be statically analyzed (like dynamic array sizes).
2.2 CPU-Side Function (Zynq ARM Core)
This function runs on Zynq’s Cortex-A9 CPU, handling all control and data preparation/processing:
#include "xil_printf.h" #include "fpga_signal_filter.h" // Auto-generated by Vivado HLS when exporting IP // CPU controller function: runs on Zynq's ARM core void cpu_signal_processor() { raw_data_t input_buffer[1024]; filtered_data_t output_buffer[1024]; int sample_count = 1024; bool enable_pipeline = true; // Step 1: Initialize the FPGA accelerator IP (using Xilinx SDK APIs) Fpga_signal_filter accelerator; fpga_signal_filter_initialize(&accelerator); // Step 2: Prepare input data (example: read from ADC or file) for (int i = 0; i < sample_count; i++) { input_buffer[i] = (i * 3) % 256; // Dummy raw sensor data } // Step 3: Configure and trigger the FPGA accelerator fpga_signal_filter_set_input_samples(&accelerator, input_buffer); fpga_signal_filter_set_sample_count(&accelerator, sample_count); fpga_signal_filter_set_enable_pipeline(&accelerator, enable_pipeline); fpga_signal_filter_start(&accelerator); // Step 4: Wait for computation to finish and read results while (!fpga_signal_filter_is_done(&accelerator)); fpga_signal_filter_get_output_samples(&accelerator, output_buffer); // Step 5: Process results (example: print to UART or send to network) xil_printf("Filtered Signal Samples (First 10):\n"); for (int i = 0; i < 10; i++) { xil_printf("Sample %d: %f\n", i, (float)output_buffer[i]); } // Step 6: Cleanup fpga_signal_filter_finalize(&accelerator); }
Key Notes for CPU-Side Code:
- Use the auto-generated header from Vivado HLS—this abstracts the low-level AXI communication into simple function calls.
- All control logic (like error handling, data sourcing) lives here; the FPGA only does what it’s told.
- Test HLS Compatibility Early: Run Vivado HLS’s "C Synthesis" on your FPGA function before integrating with the CPU code—this catches non-synthesizable code (like
mallocor function pointers) quickly. - Optimize Data Transfers: Use burst transfers (enabled via
m_axiinterfaces) instead of single-word transfers to maximize bandwidth between CPU and FPGA. - Match Bus Widths: Ensure your data types align with Zynq’s AXI bus width (32 or 64 bits) to avoid unnecessary data padding.
内容的提问来源于stack exchange,提问作者MahD

