Rust屏幕捕获库:最大化窗口捕获异常排查与解决
Rust屏幕捕获库最大化窗口异常:解决DXGI纹理行对齐(RowPitch)问题
我用Rust开发了一个简易屏幕捕获库,捕获显示器和非最大化窗口时都正常,但捕获最大化窗口时出现异常。代码大部分参考OBS的winrt-capture.cpp,但初期无法定位问题原因。
核心代码片段
// Create Frame Pool let frame_pool = Direct3D11CaptureFramePool::Create( &direct3d_device, DirectXPixelFormat::R8G8B8A8UIntNormalized, 1, item.Size()?, )?; // Create Capture Session let session = frame_pool.CreateCaptureSession(&item)?; // On Frame Arrived frame_pool.FrameArrived( &TypedEventHandler::<Direct3D11CaptureFramePool, IInspectable>::new({ move |frame, _| { // Get Texture Stuff let frame = frame.as_ref().unwrap().TryGetNextFrame()?; let frame_content_size = frame.ContentSize()?; let frame_surface = frame.Surface()?; let frame_surface = frame_surface.cast::<IDirect3DDxgiInterfaceAccess>()?; let frame_surface = unsafe { frame_surface.GetInterface::<ID3D11Texture2D>()? }; let mut desc = D3D11_TEXTURE2D_DESC::default(); unsafe { frame_surface.GetDesc(&mut desc) }; if desc.Format == DXGI_FORMAT_R8G8B8A8_UNORM { if frame_content_size.Width != last_size.Width || frame_content_size.Height != last_size.Height { let direct3d_device_recreate = &direct3d_device_recreate; frame_pool_recreate .Recreate( &direct3d_device_recreate.inner, DirectXPixelFormat::R8G8B8A8UIntNormalized, 1, frame_content_size, ) .unwrap(); last_size = frame_content_size; return Ok(()); } let texture_width = desc.Width; let texture_height = desc.Height; // Copy Texture Settings let texture_desc = D3D11_TEXTURE2D_DESC { Width: texture_width, Height: texture_height, MipLevels: 1, ArraySize: 1, Format: DXGI_FORMAT_R8G8B8A8_UNORM, SampleDesc: DXGI_SAMPLE_DESC { Count: 1, Quality: 0, }, Usage: D3D11_USAGE_STAGING, BindFlags: 0, CPUAccessFlags: D3D11_CPU_ACCESS_READ.0 as u32, MiscFlags: 0, }; // Create A Texture That CPU Can Read let mut texture = None; unsafe { d3d_device_frame_pool.CreateTexture2D( &texture_desc, None, Some(&mut texture), )? }; let texture = texture.unwrap(); // Copy The Real Texture To Copy Texture unsafe { context.CopyResource(&texture, &frame_surface) }; // Map The Texture To Enable CPU Access let mut mapped_resource = D3D11_MAPPED_SUBRESOURCE::default(); unsafe { context.Map( &texture, 0, D3D11_MAP_READ, 0, Some(&mut mapped_resource), )? }; // Create A Slice From The Bits let slice: &[Rgba] = unsafe { std::slice::from_raw_parts( mapped_resource.pData as *const Rgba, (texture_desc.Height * mapped_resource.RowPitch) as usize / std::mem::size_of::<Rgba>(), ) }; // Send The Frame To Callback Struct trigger_frame_pool.lock().on_frame_arrived( slice, texture_width, texture_height, ); // Unmap Copy Texture unsafe { context.Unmap(&texture, 0) }; } Result::Ok(()) } }), )?;
排查过程
- 发现帧池尺寸为1922x1033(显示器分辨率1920x1080),但OBS的item尺寸同样是1922x1033,排除该因素。
- 编辑1:进一步排查发现部分像素alpha值不为255,推测读取映射资源时出错。查阅DXGI文档得知,运行时可能在行之间添加padding,导致
RowPitch值超出预期。 - 编辑2:验证猜想:
窗口尺寸:1922x1033
预期字节数:1922 * 1033 * 4 = 7941704
实际字节数:Height * mapped_resource.RowPitch = 8065664
确认存在padding问题。 - 编辑3:通过逐行拷贝的方式修复问题,代码如下:
let mut vec = Vec::new(); let slice = if texture_desc.Width * texture_desc.Height * 4 == texture_desc.Height * mapped_resource.RowPitch { // 无padding,直接使用原始数据 unsafe { std::slice::from_raw_parts( mapped_resource.pData as *const Rgba, (texture_desc.Height * mapped_resource.RowPitch) as usize / std::mem::size_of::<Rgba>(), ) } } else { for i in 0..texture_desc.Height { let slice = unsafe { std::slice::from_raw_parts( mapped_resource .pData .add((i * mapped_resource.RowPitch) as usize) as *mut Rgba, texture_desc.Width as usize, ) }; vec.extend_from_slice(slice); } vec.as_slice() };
该方案可正常运行,但不确定是否为最优实现,寻求更高效的解决方法。
优化解决方法
1. 高效内存拷贝替代extend_from_slice
extend_from_slice会有额外的边界检查,直接使用底层指针拷贝可以提升效率。预分配足够容量的Vec,然后用std::ptr::copy_nonoverlapping批量拷贝每行数据:
// 预分配足够容量,避免动态扩容 let mut vec = Vec::with_capacity((texture_desc.Width * texture_desc.Height) as usize); unsafe { // 直接设置长度,跳过初始化(因为马上要覆盖数据) vec.set_len((texture_desc.Width * texture_desc.Height) as usize); for i in 0..texture_desc.Height { let src_ptr = mapped_resource.pData.add((i * mapped_resource.RowPitch) as usize) as *const Rgba; let dst_ptr = vec.as_mut_ptr().add((i * texture_desc.Width) as usize); // 批量拷贝一行的像素数据 std::ptr::copy_nonoverlapping(src_ptr, dst_ptr, texture_desc.Width as usize); } } let slice = vec.as_slice();
2. 尝试让Staging纹理使用无对齐布局(可选)
DXGI的纹理行对齐通常是4KB(4096字节),可以尝试在创建staging纹理时手动指定行字节数,但注意DXGI可能会强制覆盖这个值,所以可靠性有限:
let mut texture_desc = D3D11_TEXTURE2D_DESC { Width: texture_width, Height: texture_height, MipLevels: 1, ArraySize: 1, Format: DXGI_FORMAT_R8G8B8A8_UNORM, SampleDesc: DXGI_SAMPLE_DESC { Count: 1, Quality: 0 }, Usage: D3D11_USAGE_STAGING, BindFlags: 0, CPUAccessFlags: D3D11_CPU_ACCESS_READ.0 as u32, MiscFlags: 0, }; // 计算期望的行字节数(无padding) let desired_row_bytes = texture_desc.Width * std::mem::size_of::<Rgba>(); // 注意:DXGI可能忽略这个设置,因为staging纹理的行对齐由设备决定
3. 优化Map参数(实时场景可选)
如果是实时捕获场景,可以在Map时添加D3D11_MAP_FLAG_DO_NOT_WAIT标志,避免等待GPU,提升捕获流畅度:
unsafe { context.Map( &texture, 0, D3D11_MAP_READ, D3D11_MAP_FLAG_DO_NOT_WAIT.0 as u32, Some(&mut mapped_resource), )? };
需要确保GPU已经完成纹理拷贝操作,否则会返回错误,适合有帧同步逻辑的场景。
内容的提问来源于stack exchange,提问作者NightmareXD
相关产品推荐
相关产品推荐

