You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用_mm_cmpeq_epi8定位m128i字符向量中目标字符后,如何高效检测其下一个字节?

Efficiently Detect Next Character After Target Positions in an m128i Vector

Awesome question! Since you’ve already got the match mask from _mm_cmpeq_epi8() for your target character, you can efficiently detect the following character in the same vector with just a couple of SSE intrinsics—no need to re-scan the data all over again. Here's how to do it:

Core Idea

The key trick is to shift your existing match mask to align with the position of the next character, then combine it with a new match mask for the character you want to check. This lets you pinpoint positions where:

  • The previous byte is your original target character, AND
  • The current byte is your desired "next" character.

Step-by-Step Implementation

Let’s break this down with practical code (using x86 SSE intrinsics in C/C++):

  • Start with your existing data and match mask
    Assume you already have these values set up:

    __m128i data_vec = ...;       // Your input character data (16x 8-bit elements)
    __m128i target_char = _mm_set1_epi8('A');  // Your original target character
    __m128i target_match_mask = _mm_cmpeq_epi8(data_vec, target_char);  // Mask: 0xFF where bytes == 'A', 0x00 otherwise
    
  • Shift the match mask to target next positions
    Use _mm_srli_si128 to right-shift the mask by 1 byte. This moves each match marker to the position of the next character in the vector. Empty slots (like the first byte, which has no preceding character) get filled with 0, so we don’t get false positives here:

    __m128i prev_target_mask = _mm_srli_si128(target_match_mask, 1);
    
  • Create a match mask for your "next" character
    Generate a mask for the character you want to find immediately after your target:

    __m128i next_char = _mm_set1_epi8('B');  // The character to detect right after 'A'
    __m128i next_match_mask = _mm_cmpeq_epi8(data_vec, next_char);
    
  • Combine masks to find valid positions
    Use _mm_and_si128 to intersect the two masks. The result will have 0xFF only where both conditions are satisfied (current byte is 'B', previous byte is 'A'):

    __m128i valid_positions_mask = _mm_and_si128(prev_target_mask, next_match_mask);
    

Working with the Result

  • Check for any matches: Use _mm_movemask_epi8(valid_positions_mask)—if the returned integer is non-zero, there’s at least one valid position in the vector.
  • Count valid positions: Pass the movemask result to __builtin_popcount (or your compiler’s equivalent). Each 1 in the movemask corresponds to a valid byte position:
    int valid_count = __builtin_popcount(_mm_movemask_epi8(valid_positions_mask));
    
  • Extract positions: If you need the exact indices, you can iterate over the mask or use additional intrinsics (but for most use cases, the count or existence check is sufficient).

Edge Case Notes

  • The first byte of the vector is automatically excluded (since it has no preceding character), which is handled correctly by the shift operation (the first byte of prev_target_mask is 0).
  • If your target character is in the last byte of the vector, the shift operation drops that match (since there’s no next byte in the vector)—this is correct for in-vector checks. For cross-vector checks, you’d need to carry over the last byte’s match state to the next iteration.

内容的提问来源于stack exchange,提问作者niXman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.27 13:42:31