Scikit-Learn保存特征名称至JSON文件时遭遇ndarray无法JSON序列化的TypeError错误
Scikit-Learn保存特征名称至JSON文件时遭遇ndarray无法JSON序列化的TypeError错误
我看你遇到的是个很典型的小坑——JSON序列化天生不支持NumPy的ndarray类型,这也是用Scikit-Learn处理特征名称时经常会碰到的问题。先帮你理清问题根源,再给你两种简单的解决办法:
问题根源
你代码里的terms_unigram、terms_bigram、terms_trigram都是NumPy的ndarray对象,但Python内置的json.dump()只能序列化原生数据类型(比如列表、字典、字符串、数值等),ndarray不在它的支持范围内,所以才会抛出TypeError: Object of type ndarray is not JSON serializable这个错误。
修复方案
核心思路就是把ndarray转换成Python原生的列表,有两种轻松实现的方式:
方式1:修改函数自动适配类型
在get_dict_feature_name函数里加入类型检查,自动把ndarray转成列表,这样后续调用不用额外处理:
import json import numpy as np def get_dict_feature_name(terms): feature_name = {} # 检查输入是否为ndarray,是则转为原生列表 if isinstance(terms, np.ndarray): feature_name['feature'] = terms.tolist() else: feature_name['feature'] = terms return feature_name def save_to_json(Data, json_filename): with open(json_filename, mode='w') as json_config: json.dump(Data, json_config) # 保存特征名称 save_to_json(get_dict_feature_name(terms_unigram),"tfidf_feature_name_unigram.json") save_to_json(get_dict_feature_name(terms_bigram),"tfidf_feature_name_bigram.json") save_to_json(get_dict_feature_name(terms_trigram),"tfidf_feature_name_trigram.json")
方式2:调用时直接转换
如果不想修改原有函数,也可以在传入参数的时候,直接用.tolist()把ndarray转成列表:
# 保存特征名称 save_to_json(get_dict_feature_name(terms_unigram.tolist()),"tfidf_feature_name_unigram.json") save_to_json(get_dict_feature_name(terms_bigram.tolist()),"tfidf_feature_name_bigram.json") save_to_json(get_dict_feature_name(terms_trigram.tolist()),"tfidf_feature_name_trigram.json")
补充说明
你提供的示例三元组数据看起来是TF-IDF输出的特征权重值,转成列表后就能正常被JSON序列化,写入文件后的格式和普通Python列表一致,后续读取也能直接用json.load()解析。
备注:内容来源于stack exchange,提问作者research
相关产品推荐
相关产品推荐

