You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

HTML转PDF分页:长内容拆分至Subsection的代码修复求助

PDF内容分段处理问题修复方案

问题背景

基于HTML制作PDF页面,已完成页面设计与章节格式化,当前需处理sections数组的内容拆分:当subsection中details的content字符长度超过MAX_CHARS_PER_PAGE限制时,需将内容截断并生成新的subsection承载剩余内容,但现有TypeScript代码无法实现预期效果,仅能生成截断上限长度的subSections数组。

原始数据结构示例

const sectionsDemo = [
  {
    sectionPrefix: 'A',
    sectionHeading: 'Section A: Description of Project Activity',
    subSections: [
      {
        heading: 'Purpose and General description of project activity',
        details: [
          {
            title: 'test A1 title',
            content: 'test A1 content',
          },
        ],
      },
    ],
  },
  // ... 更多section项
];

要求的处理逻辑

  • 遍历每个section
  • 遍历section下的subSections
  • 遍历每个subsection的details数组
  • 计算details中content的字符长度
  • 若content长度超过MAX_CHARS_PER_PAGE限制,将当前content截断至上限长度
  • 将截断后的剩余内容作为新的details对象
  • 复制原subsection,携带包含剩余内容的details数组,生成新的subsection项插入到原subsection之后

期望输出示例

const sectionsDemo = [
  {
    sectionPrefix: 'A',
    sectionHeading: 'Section A: Description of Project Activity',
    subSections: [
      {
        heading: 'Purpose and General description of project activity',
        details: [
          {
            title: 'test A1 title',
            content: 'test A1 ', // 截断至上限长度的内容
          },
        ],
      },
      {
        heading: 'Purpose and General description of project activity',
        details: [
          {
            title: 'test A1 title',
            content: 'content', // 剩余内容
          },
        ],
      },
    ],
  },
  // ... 更多section项
];

现有问题代码

sectionFormatter = (sections: any) => {
  let formattedSection: any = [...sections];
  formattedSection = sections.map((section: any) => {
    return {
      ...section,
      subSections: subSectionsFormatter(section.subSections)
    }
  })
  return formattedSection
}

subSectionsFormatter = (data: any) => {
  // ** [subSections] ->[{ [details]}]-> {content, title}

  const subSections = [...data] //subsections array

  let formattedSubSections2: any = []

  for (let i = 0; i < subSections.length; i++) {
  
      const detailsArr = subSections[i].details
      if (detailsArr?.length){
        let CPCC = 0 // current page character count
        for (let j = 0; j < detailsArr.length; j++) {
          console.log('subSections[i] 44', subSections[i], data)
          const content = detailsArr[j].content
          if (content) {
            const CICC = content?.length || 0; // current item character count
             
            CPCC += CICC // sum of present chars and previously added chars
            if (CPCC > MAX_CHARS_PER_PAGE && content) {
              const LEFT_CHARS = CPCC - MAX_CHARS_PER_PAGE
              
              detailsArr[j].content = content?.slice(0, CICC - LEFT_CHARS)
              
              const newSubSection = {
                ...subSections[i],
                details:[...detailsArr.slice(0, j),{...subSections[i].details[j], content: content?.slice(CICC - LEFT_CHARS)} ]
              } 
            
              const arr: any = [subSections[i]].concat(([...subSections.slice(0, i ), newSubSection]))
             
              formattedSubSections2 = (formattedSubSections2.concat([...arr]))
              CPCC = 0
              break
            }
             else {
              formattedSubSections2.push(subSections[i])
            }
          }
        }}

     
    }

  return formattedSubSections2
}

问题分析与修复方案

现有代码的核心问题

  1. 逻辑混乱:处理剩余内容时错误拼接原数组切片,导致新subsection插入位置错误
  2. 未处理多段拆分:仅处理一次截断,剩余内容仍超限时无法继续拆分
  3. 重复添加原subsection:else分支直接push原subsection,导致重复项
  4. 变量复用错误:直接修改原detailsArr内容,污染原始数据

修正后的代码

// 定义每页最大字符数,可根据实际需求调整
const MAX_CHARS_PER_PAGE = 10;

// 拆分单个detail的content为符合长度要求的多个details对象
const splitDetailContent = (detail: {title: string, content: string}): Array<{title: string, content: string}> => {
  const result = [];
  let remainingContent = detail.content;
  
  while (remainingContent.length > MAX_CHARS_PER_PAGE) {
    // 截取当前页内容
    const pageContent = remainingContent.slice(0, MAX_CHARS_PER_PAGE);
    result.push({...detail, content: pageContent});
    // 更新剩余内容
    remainingContent = remainingContent.slice(MAX_CHARS_PER_PAGE);
  }
  
  // 添加最后一段剩余内容
  if (remainingContent.length > 0) {
    result.push({...detail, content: remainingContent});
  }
  
  return result;
};

// 处理单个subsection,拆分其details并生成对应多个subsection
const processSingleSubSection = (subSection: any): Array<any> => {
  let splitDetails: Array<{title: string, content: string}> = [];
  
  // 遍历所有detail,拆分内容
  for (const detail of subSection.details) {
    splitDetails.push(...splitDetailContent(detail));
  }
  
  // 无需拆分则直接返回原subsection
  if (splitDetails.length === subSection.details.length) {
    return [subSection];
  }
  
  // 按页面字符上限组合拆分后的details,生成多个subsection
  const subSections = [];
  let currentPageChars = 0;
  let currentDetails: Array<{title: string, content: string}> = [];
  
  for (const splitDetail of splitDetails) {
    const contentLength = splitDetail.content.length;
    
    // 当前页容量不足时,生成新subsection
    if (currentPageChars + contentLength > MAX_CHARS_PER_PAGE && currentDetails.length > 0) {
      subSections.push({...subSection, details: [...currentDetails]});
      currentDetails = [];
      currentPageChars = 0;
    }
    
    currentDetails.push(splitDetail);
    currentPageChars += contentLength;
  }
  
  // 添加最后一个未完成的subsection
  if (currentDetails.length > 0) {
    subSections.push({...subSection, details: [...currentDetails]});
  }
  
  return subSections;
};

const subSectionsFormatter = (data: any): Array<any> => {
  const formatted = [];
  for (const subSection of data) {
    formatted.push(...processSingleSubSection(subSection));
  }
  return formatted;
};

const sectionFormatter = (sections: any): Array<any> => {
  return sections.map(section => ({
    ...section,
    subSections: subSectionsFormatter(section.subSections)
  }));
};

代码说明

  1. splitDetailContent:专注处理单个detail的长内容拆分,支持多次截断直到剩余内容符合长度要求
  2. processSingleSubSection:先拆分当前subsection的所有details,再按页面字符上限将拆分后的details重新组合为多个subsection
  3. subSectionsFormatter:遍历所有原始subsection,收集处理后的结果
  4. sectionFormatter:遍历所有section,替换其subSections为处理后的数组

该代码支持多层拆分、不污染原始数据,完全符合需求的处理逻辑。


内容的提问来源于stack exchange,提问作者Saumay Paul

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 12:48:08