You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

HEVC参考软件修改:按行编码CTU实现波前编码的方案咨询

关于HM编码器实现波前(按行)编码的方案解析

Great question—let’s unpack this clearly based on how HM’s encoding pipeline works, especially around wavefront parallel processing (WPP).

First, let’s get straight to the point: Setting the slice segment to the full frame width is NOT the right approach to implement row-based wavefront encoding. Here’s why:

1. How the current slice-based CTU encoding works

As you observed, the for( UInt ctuTsAddr = startCtuTsAddr; ctuTsAddr < boundingCtuTsAddr; ++ctuTsAddr ) loop processes CTUs in raster scan order within a single slice. Slices are independent encoding units by design, so all CTUs inside a slice are processed sequentially by default. Even if you stretch a slice to cover an entire row of CTUs, those CTUs will still be encoded one after another in order—no parallelism here, which defeats the core goal of wavefront encoding.

2. The correct way to enable wavefront (row-based) parallel encoding

HM already has built-in support for WPP, which is purpose-built for row-based parallel processing of CTUs. You don’t need to tweak slice sizes to make this work:

  • First, enable WPP in your configuration file with WavefrontParallelProcessing=1
  • Under the hood, HM’s encoder maintains a dependency tracking system for CTUs across rows. It allows a CTU in row N to start encoding as soon as the CTUs in row N-1 (usually 1-2 positions to the left) have finished. This creates the "wavefront" of parallel processing across rows.
  • This logic operates at the CTU scheduling layer, not the slice layer. You can keep your slices at their original size (even full-frame slices) and still reap WPP’s parallel benefits.

3. What happens if you force slices to be full-row width?

Doing this would actually backfire:

  • You’d lose any slice-level parallelism you might have had before
  • The row’s CTUs would still process sequentially (thanks to that slice inner loop), so you gain no wavefront-style parallelism
  • You’d also introduce unnecessary constraints on intra prediction and entropy coding, since slices enforce independent context resetting which isn’t needed for WPP.

If you’re aiming to implement a custom row-based encoding logic (instead of using HM’s built-in WPP), you’d need to modify the CTU scheduling logic directly—not slice sizes. For example, you could set up worker threads that handle individual rows, with synchronization to ensure a row’s CTUs only start after the required upstream CTUs from the row above are completed.


内容的提问来源于stack exchange,提问作者Maverick

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 07:14:37