LLM token限制下大体积HTML文档分块与修改方案咨询
CI/CD流水线中基于LLM的HTML批量修改方案优化
场景
在CI/CD流水线新增处理层,部署前根据用户/开发者反馈修改项目HTML构建文件,比如要求“确保图片懒加载”这类优化需求。
当前问题
- 处理的HTML文档均至少包含248k token,远超LLM上下文窗口(128k),无法直接传入完整文档
- 已设置单块token上限15k(预留空间给prompt,LLM输出约16k),但现有深度优先分块脚本存在块间重复元素
- 尝试让LLM仅返回关联修改,但LLM需要完整HTML树结构才能准确判断,因此分块策略是核心痛点
需求目标
- 脚本将HTML分块为最大15k token、无重复内容的块
- 修改后的块可直接拼接为完整的有效HTML文件
- 欢迎探索更优的替代方案
原分块脚本(存在重复元素问题)
def split_html_into_chunks(self, node, token_limit): chunks = [] current_chunk = [] current_tokens = 0 def dfs(element): nonlocal current_chunk, current_tokens # Get text and token count of the element if isinstance(element, str): text = element.strip() tokens = self.estimate_tokens(text) else: text = str(element) tokens = self.estimate_tokens(text) # If a single element exceeds the token limit, break it down further if tokens > token_limit: if hasattr(element, 'children'): for child in element.children: dfs(child) return # If adding this element exceeds the limit, finalize the current chunk if current_tokens + tokens > token_limit: chunks.append(current_chunk) current_chunk = [] current_tokens = 0 # Add element to current chunk current_chunk.append(element) current_tokens += tokens # Recurse into children if it's a Tag if hasattr(element, 'children'): for child in element.children: dfs(child) # Start DFS traversal dfs(node) # Add the last chunk if not empty if current_chunk: chunks.append(current_chunk) return chunks
原脚本问题分析
- 递归处理子元素时,父元素已被加入当前块,子元素又被重复加入,导致块间出现重复标签/文本
- 分块时未区分“完整标签单元”,可能把单个标签的开闭合拆分到不同块中
修复后的无重复分块脚本
from bs4 import BeautifulSoup, Tag, NavigableString def split_html_into_chunks(self, root_node, token_limit): chunks = [] current_chunk_elements = [] current_token_count = 0 def process_element(element): nonlocal current_chunk_elements, current_token_count # 计算当前元素的token数(包含完整标签结构) if isinstance(element, NavigableString): element_str = str(element).strip() if not element_str: return # 跳过空白文本 token_count = self.estimate_tokens(element_str) elif isinstance(element, Tag): element_str = str(element) token_count = self.estimate_tokens(element_str) else: return # 跳过非标签/文本元素 # 如果单个元素超过token限制,递归拆分其子元素(仅针对标签) if token_count > token_limit: if isinstance(element, Tag) and element.children: for child in element.children: process_element(child) return # 检查加入当前块是否超限 if current_token_count + token_count > token_limit: # 保存当前块并重置 chunks.append(current_chunk_elements.copy()) current_chunk_elements = [] current_token_count = 0 # 添加元素到当前块 current_chunk_elements.append(element) current_token_count += token_count # 从根节点开始处理 process_element(root_node) # 加入最后一个非空块 if current_chunk_elements: chunks.append(current_chunk_elements) return chunks # 序列化块为可拼接的HTML字符串 def serialize_chunks(self, chunks): serialized = [] for chunk in chunks: chunk_str = ''.join(str(elem) for elem in chunk) serialized.append(chunk_str) return serialized
修复说明
- 仅在元素完整时加入块,避免拆分标签结构
- 跳过空白文本节点,减少无效token占用
- 递归拆分大元素时,不再将父元素加入块,仅处理子元素,避免重复
- 确保每个块都是独立无重复的HTML片段,拼接后可恢复完整文档
优化后的修改与拼接脚本
def modify_html_with_llm(self, audit_issue): soup = BeautifulSoup(self.html, 'html.parser') # 分块(使用修复后的脚本) html_chunks = self.split_html_into_chunks(soup, TOKEN_LIMIT) serialized_chunks = self.serialize_chunks(html_chunks) modified_chunks = [] for i, chunk in enumerate(serialized_chunks): print(f"处理第 {i+1}/{len(serialized_chunks)} 块,token数:{self.estimate_tokens(chunk)}") modified_chunk = self.send_to_llm(chunk, audit_issue) modified_chunks.append(modified_chunk) # 直接拼接即可得到完整修改后的HTML modified_html = ''.join(modified_chunks) # 可选:验证HTML有效性 try: BeautifulSoup(modified_html, 'html.parser') except Exception as e: print(f"拼接后HTML无效:{e}") return self.html # 回退原HTML return modified_html
优化后的LLM提示词
你是资深前端工程师,根据以下优化需求:{feedback},修改给定的HTML片段。 要求: 1. 仅针对需求做必要修改,不改动无关内容 2. 必须返回完整的HTML片段,保持原有的标签结构完整性 3. 如果无法针对该片段进行修改,直接返回原片段 4. 不要添加任何额外解释或说明,只返回HTML代码 待修改的HTML片段: {html_part}
提示词优化点
- 明确要求保持标签结构完整性,避免LLM返回不完整的标签
- 强调仅返回HTML代码,减少无关输出
更优替代方案:基于语义的分块策略
对于大体积HTML,可优先按语义单元分块(如<header>、<main>、<section>等),再对超大语义单元进行细拆分:
def split_html_by_semantic_units(self, soup, token_limit): semantic_tags = ['header', 'nav', 'main', 'section', 'article', 'aside', 'footer'] chunks = [] # 先按语义标签拆分大单元 for tag_name in semantic_tags: elements = soup.find_all(tag_name) for elem in elements: elem_str = str(elem) elem_tokens = self.estimate_tokens(elem_str) if elem_tokens <= token_limit: chunks.append([elem]) else: # 对超大语义单元进行递归拆分 sub_chunks = self.split_html_into_chunks(elem, token_limit) chunks.extend(sub_chunks) # 处理剩余非语义标签内容 remaining_content = [] for child in soup.children: if not isinstance(child, Tag) or child.name not in semantic_tags: remaining_content.append(child) if remaining_content: remaining_str = ''.join(str(c) for c in remaining_content) if self.estimate_tokens(remaining_str) <= token_limit: chunks.append(remaining_content) else: sub_chunks = self.split_html_into_chunks(BeautifulSoup(remaining_str, 'html.parser'), token_limit) chunks.extend(sub_chunks) return chunks
优势
- 语义单元分块更符合LLM的理解逻辑,更容易定位需要修改的部分(比如图片懒加载主要在
<main>或<section>中) - 减少不必要的分块数量,提升处理效率
内容的提问来源于stack exchange,提问作者Pogo
相关产品推荐
相关产品推荐

