You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按指定属性名解析文本键值对并转换写入JSON文件

问题描述

我正在编写程序生成JSON文件,存储从指定格式文本中匹配提取到的字段内容,待处理的文本格式如下:

servicepoint ‏ 200135644 watchid ‏ 7038842

需求是将servicepoint、watchid的对应值仅插入对象的table数组一次,当前已编写的基础实现代码如下:

function readfile() {
  Tesseract.recognize('form.png', 'ara', {

    logger: m => console.log(m)

  }).then(({ data: { text } }) => {
    console.log(text); /* this line here */
    var obj = {
      table: []
    };
    const info = ['servicepoint', 'watchid'];
    for (k = 0; k < info.length; k++) {
      var result = text.match(new RegExp(info[k] + '\\s+(\\w+)'))[1];
      obj.table.push({
        servicepoint: /* Here i want to insert the number after servicepoint*/ ,
        watchid: /*also i want to insert the number after watchid to the Object table*/
      });
    }
    var json = JSON.stringify(obj); /* converting the object table to json file*/
    var fs = require('fs'); /* and then write json file contians the data*/
    fs.writeFile('myjsonfile.json', json, 'utf8', callback);
  })
};

当前需要补全代码逻辑:分别提取文本中servicepoint、watchid关键词后跟随的数值,填入对应对象属性位,最终将构造完成的对象转换为JSON格式写入本地文件。

实现方案

原代码存在两个核心问题:一是通过循环遍历字段的逻辑会重复向table数组插入对象,最终生成两条冗余记录;二是使用阿拉伯语OCR模型时,识别结果很容易夹带不可见的控制字符,原正则仅匹配空白字符容易出现匹配失败的问题。
修正后的可运行代码如下:

function readfile() {
  const fs = require('fs');
  Tesseract.recognize('form.png', 'ara', {
    logger: m => console.log(m)
  }).then(({ data: { text } }) => {
    console.log('OCR识别原始文本:', text);
    const obj = { table: [] };
    // 分别匹配两个字段值,兼容不可见特殊字符
    const spMatch = text.match(/servicepoint\s*\W*(\w+)/);
    const widMatch = text.match(/watchid\s*\W*(\w+)/);

    if (!spMatch || !widMatch) {
      throw new Error('字段匹配失败,请检查图片识别结果');
    }
    // 仅插入一次记录到table数组
    obj.table.push({
      servicepoint: spMatch[1],
      watchid: widMatch[1]
    });

    const json = JSON.stringify(obj, null, 2);
    fs.writeFile('myjsonfile.json', json, 'utf8', (err) => {
      if (err) throw err;
      console.log('JSON文件写入完成');
    });
  }).catch(err => console.error('执行异常:', err));
}

主要调整点:

  • 删除原有的字段遍历循环,直接单独匹配两个目标字段,保证table数组仅写入一条包含两个字段的记录
  • 正则增加\W*匹配规则,兼容OCR带出的不可见控制字符、特殊符号,降低匹配失败概率
  • 增加匹配结果校验、异常捕获、文件写入回调逻辑,避免无提示报错
  • JSON.stringify增加缩进参数,生成的JSON文件格式更易读
  • 调整fs依赖的引入位置,无需等待OCR识别完成再加载模块

内容的提问来源于stack exchange,提问作者MrObscure

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.30 15:03:23