如何用Python计算JSON文件中分类的百分位数?
优化分类统计代码并计算分类百分位数
原代码的性能问题
你写的countCategory函数用列表存储分类,每次判断category not in categorieslist是**O(n)**时间复杂度——列表的成员检查需要遍历整个列表,处理大型JSON时会越跑越慢,数据量越大性能下降越明显。
优化分类获取代码
用集合(set)替代列表,集合的成员检查是**O(1)**时间复杂度,自动去重,还能直接用update批量添加分类,效率提升显著:
def get_all_categories(file): categories_set = set() for book in file: categories_set.update(book["categories"]) return list(categories_set)
统计分类书籍数量并计算百分位数
要计算百分位数,首先得统计每个分类对应的书籍数量,再基于数量排序计算百分位:
- 统计分类书籍数(用
collections.Counter更简洁)
from collections import Counter def count_category_books(file): category_counter = Counter() for book in file: category_counter.update(book["categories"]) return category_counter
- 计算百分位数
def calculate_category_percentiles(category_counter): # 把所有分类的书籍数量排序,用于计算排名 sorted_counts = sorted(category_counter.values()) total_categories = len(sorted_counts) percentiles = {} for category, count in category_counter.items(): # 统计有多少个分类的书籍数 ≤ 当前分类的数量 rank = sum(1 for c in sorted_counts if c <= count) # 计算百分位数(保留两位小数) percentile = (rank / total_categories) * 100 percentiles[category] = round(percentile, 2) return percentiles
处理大型JSON的注意事项
如果JSON文件特别大,别一次性把整个文件加载到内存,建议用ijson库流式读取,避免内存溢出:
import ijson def load_large_json(file_path): with open(file_path, 'r', encoding='utf-8') as f: # 流式读取每个book对象 yield from ijson.items(f, 'item')
主流程示例
if __name__ == "__main__": # 假设你已经把JSON文件下载到本地 books = load_large_json("books.json") category_counts = count_category_books(books) category_percentiles = calculate_category_percentiles(category_counts) # 打印结果 for category, p in category_percentiles.items(): print(f"分类「{category}」: 百分位数 {p}%")
内容的提问来源于stack exchange,提问作者Miguel Angelo
相关产品推荐
相关产品推荐

