Python调用seoanalyzer如何将分析结果正确写入JSON文件
问题根因
- 现有代码只在文件操作的上下文块中调用
print(output),仅将结果输出到终端,未执行任何写入JSON文件的逻辑 - seoanalyzer返回的分析结果中,
bigrams等词频统计字段是collections.Counter类型,Python内置JSON模块默认不支持该类型的序列化,直接写入会触发类型报错
修复方案
核心处理逻辑分两步:一是调用标准库json模块的dump方法执行文件写入,二是提前把结果中所有Counter类型的对象转为原生字典,解决序列化兼容问题。
可直接运行的完整代码如下:
from seoanalyzer import analyze import json from collections import Counter site = 'http://www.microblink.com' output = analyze(site, follow_links=False) # 递归遍历结果,将所有Counter对象转为普通字典,适配JSON序列化要求 def parse_counter(data): if isinstance(data, Counter): return dict(data) if isinstance(data, list): return [parse_counter(item) for item in data] if isinstance(data, dict): return {key: parse_counter(value) for key, value in data.items()} return data parsed_result = parse_counter(output) with open("output.json", "w", encoding='utf-8') as file: # 写入时保留缩进格式化,关闭ASCII转义保证特殊字符正常存储 json.dump(parsed_result, file, ensure_ascii=False, indent=2) # 如需同时在终端打印分析结果,保留该行即可 print(parsed_result)
说明
- 用到的
json、collections.Counter都是Python标准库内置模块,不需要额外安装第三方包 - 类型转换不会丢失任何词频统计数据,仅把Counter子类转为JSON可识别的原生字典结构
- 写入参数
indent=2会让输出的JSON文件有清晰的层级缩进,方便后续阅读解析;ensure_ascii=False可以避免多语言字符被转义为乱码
内容的提问来源于stack exchange,提问作者Tori Elstrom
相关产品推荐
相关产品推荐

