You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

LLM token限制下大体积HTML文档分块与修改方案咨询

CI/CD流水线中基于LLM的HTML批量修改方案优化

场景

在CI/CD流水线新增处理层,部署前根据用户/开发者反馈修改项目HTML构建文件,比如要求“确保图片懒加载”这类优化需求。

当前问题

  • 处理的HTML文档均至少包含248k token,远超LLM上下文窗口(128k),无法直接传入完整文档
  • 已设置单块token上限15k(预留空间给prompt,LLM输出约16k),但现有深度优先分块脚本存在块间重复元素
  • 尝试让LLM仅返回关联修改,但LLM需要完整HTML树结构才能准确判断,因此分块策略是核心痛点

需求目标

  1. 脚本将HTML分块为最大15k token、无重复内容的块
  2. 修改后的块可直接拼接为完整的有效HTML文件
  3. 欢迎探索更优的替代方案

原分块脚本(存在重复元素问题)

def split_html_into_chunks(self, node, token_limit):
        chunks = []
        current_chunk = []
        current_tokens = 0

        def dfs(element):
            nonlocal current_chunk, current_tokens

            # Get text and token count of the element
            if isinstance(element, str):
                text = element.strip()
                tokens = self.estimate_tokens(text)
            else:
                text = str(element)
                tokens = self.estimate_tokens(text)

            # If a single element exceeds the token limit, break it down further
            if tokens > token_limit:
                if hasattr(element, 'children'):
                    for child in element.children:
                        dfs(child)
                return

            # If adding this element exceeds the limit, finalize the current chunk
            if current_tokens + tokens > token_limit:
                chunks.append(current_chunk)
                current_chunk = []
                current_tokens = 0

            # Add element to current chunk
            current_chunk.append(element)
            current_tokens += tokens

            # Recurse into children if it's a Tag
            if hasattr(element, 'children'):
                for child in element.children:
                    dfs(child)

        # Start DFS traversal
        dfs(node)

        # Add the last chunk if not empty
        if current_chunk:
            chunks.append(current_chunk)

        return chunks

原脚本问题分析

  • 递归处理子元素时,父元素已被加入当前块,子元素又被重复加入,导致块间出现重复标签/文本
  • 分块时未区分“完整标签单元”,可能把单个标签的开闭合拆分到不同块中

修复后的无重复分块脚本

from bs4 import BeautifulSoup, Tag, NavigableString

def split_html_into_chunks(self, root_node, token_limit):
    chunks = []
    current_chunk_elements = []
    current_token_count = 0

    def process_element(element):
        nonlocal current_chunk_elements, current_token_count

        # 计算当前元素的token数(包含完整标签结构)
        if isinstance(element, NavigableString):
            element_str = str(element).strip()
            if not element_str:
                return  # 跳过空白文本
            token_count = self.estimate_tokens(element_str)
        elif isinstance(element, Tag):
            element_str = str(element)
            token_count = self.estimate_tokens(element_str)
        else:
            return  # 跳过非标签/文本元素

        # 如果单个元素超过token限制,递归拆分其子元素(仅针对标签)
        if token_count > token_limit:
            if isinstance(element, Tag) and element.children:
                for child in element.children:
                    process_element(child)
            return

        # 检查加入当前块是否超限
        if current_token_count + token_count > token_limit:
            # 保存当前块并重置
            chunks.append(current_chunk_elements.copy())
            current_chunk_elements = []
            current_token_count = 0

        # 添加元素到当前块
        current_chunk_elements.append(element)
        current_token_count += token_count

    # 从根节点开始处理
    process_element(root_node)

    # 加入最后一个非空块
    if current_chunk_elements:
        chunks.append(current_chunk_elements)

    return chunks

# 序列化块为可拼接的HTML字符串
def serialize_chunks(self, chunks):
    serialized = []
    for chunk in chunks:
        chunk_str = ''.join(str(elem) for elem in chunk)
        serialized.append(chunk_str)
    return serialized

修复说明

  • 仅在元素完整时加入块,避免拆分标签结构
  • 跳过空白文本节点,减少无效token占用
  • 递归拆分大元素时,不再将父元素加入块,仅处理子元素,避免重复
  • 确保每个块都是独立无重复的HTML片段,拼接后可恢复完整文档

优化后的修改与拼接脚本

def modify_html_with_llm(self, audit_issue):
    soup = BeautifulSoup(self.html, 'html.parser')
    
    # 分块(使用修复后的脚本)
    html_chunks = self.split_html_into_chunks(soup, TOKEN_LIMIT)
    serialized_chunks = self.serialize_chunks(html_chunks)
  
    modified_chunks = []
    for i, chunk in enumerate(serialized_chunks):
        print(f"处理第 {i+1}/{len(serialized_chunks)} 块,token数:{self.estimate_tokens(chunk)}")
        modified_chunk = self.send_to_llm(chunk, audit_issue)
        modified_chunks.append(modified_chunk)

    # 直接拼接即可得到完整修改后的HTML
    modified_html = ''.join(modified_chunks)
    # 可选:验证HTML有效性
    try:
        BeautifulSoup(modified_html, 'html.parser')
    except Exception as e:
        print(f"拼接后HTML无效:{e}")
        return self.html  # 回退原HTML
    
    return modified_html

优化后的LLM提示词

你是资深前端工程师,根据以下优化需求:{feedback},修改给定的HTML片段。
要求:
1. 仅针对需求做必要修改,不改动无关内容
2. 必须返回完整的HTML片段,保持原有的标签结构完整性
3. 如果无法针对该片段进行修改,直接返回原片段
4. 不要添加任何额外解释或说明,只返回HTML代码

待修改的HTML片段:
{html_part}

提示词优化点

  • 明确要求保持标签结构完整性,避免LLM返回不完整的标签
  • 强调仅返回HTML代码,减少无关输出

更优替代方案:基于语义的分块策略

对于大体积HTML,可优先按语义单元分块(如<header>、<main>、<section>等),再对超大语义单元进行细拆分:

def split_html_by_semantic_units(self, soup, token_limit):
    semantic_tags = ['header', 'nav', 'main', 'section', 'article', 'aside', 'footer']
    chunks = []

    # 先按语义标签拆分大单元
    for tag_name in semantic_tags:
        elements = soup.find_all(tag_name)
        for elem in elements:
            elem_str = str(elem)
            elem_tokens = self.estimate_tokens(elem_str)
            if elem_tokens <= token_limit:
                chunks.append([elem])
            else:
                # 对超大语义单元进行递归拆分
                sub_chunks = self.split_html_into_chunks(elem, token_limit)
                chunks.extend(sub_chunks)
    
    # 处理剩余非语义标签内容
    remaining_content = []
    for child in soup.children:
        if not isinstance(child, Tag) or child.name not in semantic_tags:
            remaining_content.append(child)
    if remaining_content:
        remaining_str = ''.join(str(c) for c in remaining_content)
        if self.estimate_tokens(remaining_str) <= token_limit:
            chunks.append(remaining_content)
        else:
            sub_chunks = self.split_html_into_chunks(BeautifulSoup(remaining_str, 'html.parser'), token_limit)
            chunks.extend(sub_chunks)
    
    return chunks

优势

  • 语义单元分块更符合LLM的理解逻辑,更容易定位需要修改的部分(比如图片懒加载主要在<main>或<section>中)
  • 减少不必要的分块数量,提升处理效率

内容的提问来源于stack exchange,提问作者Pogo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 19:53:16