如何生成带特征的文本token并输出指定结构的JSON数据
文本分词转指定格式JSON实现方案
需求描述
将给定的文本列表按空格分词为带特征的token,最终输出符合指定结构的JSON数据。
输入样例
words = ['The study of aviation safety report in the aviation industry usually relies', 'The experimental results show that compared with traditional', 'Heterogeneous Aviation Safety Cases: Integrating the Formal and the Non-formal']
目标输出结构
{"sentence": [ { "indexSentence":0, "tokens": [{ "indexWord": 1, "word": "The", "len": 3 }, { "indexWord": 2, "word": "study", "len": 5}, {"indexWord": 3, "word": "of", "len": 2 }, {"indexWord": 4, "word": "aviation", "len": 8} ] }, { "indexSentence" : 1, "tokens" : [] } ]}
原有代码问题
原有代码核心错误点:
- 用句子长度作为字典键,不同句子长度一致时会出现值覆盖
- 整体数据结构和要求的JSON结构完全不匹配
- 分词后的属性命名、索引规则不符合需求
修正后实现代码
import json # 输入文本列表 words = ['The study of aviation safety report in the aviation industry usually relies', 'The experimental results show that compared with traditional', 'Heterogeneous Aviation Safety Cases: Integrating the Formal and the Non-formal'] # 构造结果结构 result = {"sentence": []} for sent_idx, sentence in enumerate(words): # 按空格拆分句子为单词 word_list = sentence.split() tokens = [] # indexWord从1开始计数 for word_idx, word in enumerate(word_list, start=1): tokens.append({ "indexWord": word_idx, "word": word, "len": len(word) }) result["sentence"].append({ "indexSentence": sent_idx, "tokens": tokens }) # 输出格式化后的JSON print(json.dumps(result, indent=2, ensure_ascii=False))
内容的提问来源于stack exchange,提问作者hammu
相关产品推荐
相关产品推荐

