You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何仅在/>位置拆分HTML为适配AWS Translate的指定长度块?

实现方案

核心逻辑要同时满足两个约束:单块UTF-8编码字节数不超过5000、仅在/>位置拆分避免截断HTML标签,具体实现如下:

def split_html_for_aws_translate(html_content, max_chunk_bytes=5000):
    chunks = []
    current_pos = 0
    html_total_len = len(html_content)
    
    while current_pos < html_total_len:
        # 先找到当前位置开始、不超过最大字节限制的最远边界
        end_candidate = current_pos
        while end_candidate < html_total_len and len(html_content[current_pos:end_candidate+1].encode('utf-8')) <= max_chunk_bytes:
            end_candidate += 1
        
        # 在允许范围内查找最后一个 /> 作为拆分点
        split_pos = html_content.rfind('/>', current_pos, end_candidate)
        
        # 兜底逻辑:未找到/>时找最后一个>,极端情况直接按最大长度切
        if split_pos == -1:
            split_pos = html_content.rfind('>', current_pos, end_candidate)
            if split_pos == -1:
                split_pos = end_candidate - 1
            else:
                split_pos += 1 # 包含>本身
        else:
            split_pos += 2 # 包含/>本身
        
        # 存入切块并更新当前指针
        chunks.append(html_content[current_pos:split_pos])
        current_pos = split_pos
    return chunks

# 调用示例
chunks = split_html_for_aws_translate(HTML_CONTENT)

注意事项

  • 直接按UTF-8编码字节数计算块大小,完全匹配AWS Translate的校验规则,比按字符数估算的方案更准确,不会出现请求超限的问题
  • 优先按/>拆分,符合不截断HTML标签的要求,兜底逻辑避免无自闭合标签的片段导致死循环
  • 如果需要更高阶的拆分要求(比如不截断配对标签<div></div>),可以引入BeautifulSoup等HTML解析库,按DOM节点结构拆分即可

内容的提问来源于stack exchange,提问作者Milano

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 07:54:04