如何仅在/>位置拆分HTML为适配AWS Translate的指定长度块?
实现方案
核心逻辑要同时满足两个约束:单块UTF-8编码字节数不超过5000、仅在/>位置拆分避免截断HTML标签,具体实现如下:
def split_html_for_aws_translate(html_content, max_chunk_bytes=5000): chunks = [] current_pos = 0 html_total_len = len(html_content) while current_pos < html_total_len: # 先找到当前位置开始、不超过最大字节限制的最远边界 end_candidate = current_pos while end_candidate < html_total_len and len(html_content[current_pos:end_candidate+1].encode('utf-8')) <= max_chunk_bytes: end_candidate += 1 # 在允许范围内查找最后一个 /> 作为拆分点 split_pos = html_content.rfind('/>', current_pos, end_candidate) # 兜底逻辑:未找到/>时找最后一个>,极端情况直接按最大长度切 if split_pos == -1: split_pos = html_content.rfind('>', current_pos, end_candidate) if split_pos == -1: split_pos = end_candidate - 1 else: split_pos += 1 # 包含>本身 else: split_pos += 2 # 包含/>本身 # 存入切块并更新当前指针 chunks.append(html_content[current_pos:split_pos]) current_pos = split_pos return chunks # 调用示例 chunks = split_html_for_aws_translate(HTML_CONTENT)
注意事项
- 直接按UTF-8编码字节数计算块大小,完全匹配AWS Translate的校验规则,比按字符数估算的方案更准确,不会出现请求超限的问题
- 优先按
/>拆分,符合不截断HTML标签的要求,兜底逻辑避免无自闭合标签的片段导致死循环 - 如果需要更高阶的拆分要求(比如不截断配对标签
<div></div>),可以引入BeautifulSoup等HTML解析库,按DOM节点结构拆分即可
内容的提问来源于stack exchange,提问作者Milano
相关产品推荐
相关产品推荐

