You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

JS实现非结构化对象数组转指定结构化Content格式

JS结构化内容数组解析方案

问题背景

我有如下JS对象数据:

const content = 
  [ { h1 : 'This is title number 1'          } 
  , { h2 : 'Description'                     }
  , { p  : 'Description unique content text' }
  , { h2 : 'Content'                         }
  , { p  : 'Content unique text here 1'      }
  , { ul : [ 
           'string value 1',
           'string value 2'
       ]                       
  , } 
  , { p  : 'Content unique text here 2'      }
  , { h2 : 'CTA message'                     }
  , { p  : 'CTA message unique content here' }
  , { h2 : 'CTA button'                      }
  , { p  : 'CTA button unique content here'  }
  , { p  : ''                                }
  , { h1 : 'This is title number 2'          } 
  , { h2 : 'Description'                     }
  , { p  : 'Description unique content text' }
  , { h2 : 'Content'                         }
  , { p  : 'Content unique text here 1'      }
  , { h2 : 'CTA message'                     }
  , { p  : 'CTA message unique content here' }
  , { h2 : 'CTA button'                      }
  , { p  : 'CTA button unique content here'  }
  , { p  : ''                                }
  ]

规则说明:

  • h1标识一个新对象的起始
  • ul为可选字段

需要将上述数据映射为符合以下TS接口定义的结构:

interface Content {
  title: string;
  description: string;
  content: any[];
  cta: {
    message: string;
    button: string;
  }
}

我最初的思路是遍历数组元素,按照接口定义填充新的JSON对象:首元素固定为title,识别到对应"Description"的h2节点后,取其下一个元素作为description字段值,初步逻辑代码如下:

const json: Content[] = [];
content.forEach(element => {
  if(element.h1) {
      // 为新对象添加属性
      // 将新对象推入json结果数组
  }
});

但我不确定如何基于原始内容数组正确生成多个Content对象,期望得到的输出结果格式如下:

[
  {
    "title": "This is title number 1",
    "description": "Description unique content text",
    "content": [
      {
        "p": "Content unique text here 1"
      },
      {
        "ul": [
            "string value 1",
            "string value 2"
        ]
      },
      {
        "p": "Content unique text here 2"
      } 
    ],
   "cta": {
      "message": "CTA message unique content here",
      "button": "CTA button unique content here"
   }
  }
]

更新要求

需要一套自上而下的解析方案,要求方案具备良好的可扩展性,当输入数组新增h2+p或h2+p+ul等组合结构时,可快速适配调整。


最优实现方案

采用状态驱动的逐行解析方案,逻辑清晰、扩展成本极低,不需要硬编码下标取数,完全适配后续结构调整:

  1. 维护当前解析状态、当前正在构建的Content对象两个核心变量
  2. 遍历原始数组时,根据当前遇到的节点类型切换解析状态,将后续节点填充到对应字段
  3. 遇到新的h1节点时,把上一个构建完成的对象存入结果数组,初始化新的当前对象

完整实现代码如下:

function parseContent(rawList) {
  const result = [];
  // 解析状态枚举,后续新增字段直接在这里加状态即可
  const ParseState = {
    INIT: 'init',
    DESCRIPTION: 'description',
    CONTENT: 'content',
    CTA_MSG: 'cta_msg',
    CTA_BTN: 'cta_btn'
  }

  let current = null;
  let state = ParseState.INIT;

  for (const item of rawList) {
    // 遇到h1,开启新的内容块
    if (item.h1) {
      if (current) result.push(current);
      current = {
        title: item.h1,
        description: '',
        content: [],
        cta: { message: '', button: '' }
      };
      state = ParseState.INIT;
      continue;
    }

    // 遇到h2,切换对应解析状态
    if (item.h2) {
      switch(item.h2) {
        case 'Description':
          state = ParseState.DESCRIPTION;
          break;
        case 'Content':
          state = ParseState.CONTENT;
          break;
        case 'CTA message':
          state = ParseState.CTA_MSG;
          break;
        case 'CTA button':
          state = ParseState.CTA_BTN;
          break;
        // 后续新增h2类型直接加case即可
        default:
          state = ParseState.CONTENT;
      }
      continue;
    }

    // 空节点直接跳过
    if (item.p === '' || !Object.keys(item).length) continue;

    // 根据当前状态填充字段
    switch(state) {
      case ParseState.DESCRIPTION:
        current.description = item.p;
        break;
      case ParseState.CTA_MSG:
        current.cta.message = item.p;
        break;
      case ParseState.CTA_BTN:
        current.cta.button = item.p;
        break;
      case ParseState.CONTENT:
        // content区块直接塞节点,不管是p还是ul都兼容,后续加ol、img等类型不用改逻辑
        current.content.push(item);
        break;
    }
  }

  // 把最后一个构建完的对象推入结果
  if (current) result.push(current);
  return result;
}

// 调用即可得到目标结果
const json = parseContent(content);

扩展说明

后续如果要新增一级字段,只需要做3件事:

  • 在ParseState里加对应状态值
  • 在h2的switch分支里加对应标题的状态映射
  • 在字段填充的switch分支里加对应赋值逻辑

如果是content区块内新增节点类型(比如图片、有序列表、视频嵌入块),完全不需要修改解析逻辑,会自动存入content数组。


内容的提问来源于stack exchange,提问作者Sergino

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.30 06:57:17