You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何生成带特征的文本token并输出指定结构的JSON数据

文本分词转指定格式JSON实现方案

需求描述

将给定的文本列表按空格分词为带特征的token,最终输出符合指定结构的JSON数据。

输入样例

words = ['The study of aviation safety report in the aviation industry usually relies', 
         'The experimental results show that compared with traditional',
         'Heterogeneous Aviation Safety Cases: Integrating the Formal and the Non-formal']

目标输出结构

{"sentence": [
           {
             "indexSentence":0,
             "tokens": [{
                       "indexWord": 1,
                        "word": "The",
                         "len": 3
                      },
                      { "indexWord": 2,
                        "word": "study",
                         "len": 5},
                      {"indexWord": 3,
                        "word": "of",
                         "len": 2
                       },
                       {"indexWord": 4,
                        "word": "aviation",
                         "len": 8}
                    ]
           },
           {
            "indexSentence" : 1,
            "tokens" : []
           }
         ]}

原有代码问题

原有代码核心错误点:

  • 用句子长度作为字典键,不同句子长度一致时会出现值覆盖
  • 整体数据结构和要求的JSON结构完全不匹配
  • 分词后的属性命名、索引规则不符合需求

修正后实现代码

import json

# 输入文本列表
words = ['The study of aviation safety report in the aviation industry usually relies', 
         'The experimental results show that compared with traditional',
         'Heterogeneous Aviation Safety Cases: Integrating the Formal and the Non-formal']

# 构造结果结构
result = {"sentence": []}
for sent_idx, sentence in enumerate(words):
    # 按空格拆分句子为单词
    word_list = sentence.split()
    tokens = []
    # indexWord从1开始计数
    for word_idx, word in enumerate(word_list, start=1):
        tokens.append({
            "indexWord": word_idx,
            "word": word,
            "len": len(word)
        })
    result["sentence"].append({
        "indexSentence": sent_idx,
        "tokens": tokens
    })

# 输出格式化后的JSON
print(json.dumps(result, indent=2, ensure_ascii=False))

内容的提问来源于stack exchange,提问作者hammu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.27 18:06:06