You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中计算JSON数据里的duration总时长

如何高效计算JSON中所有"duration"字段的总时长

我有一个结构固定但内部值动态的JSON数据,需要在Python中计算所有"duration"字段的总时长。目前通过多层嵌套for循环实现,但希望找到更简洁、通用的高效方法。以下是JSON示例及现有代码:

JSON示例

{"data": {"9": {"mp_e_b": {"english": {"english_literature": {"video_lessons": {"The Fun They Had": {"-NNH9_LnN_RBLjxN3lKN": {"class_name": "9","duration": "28","subject_name": "English Literature"},"-NNH9jblawQprJ9-ZFmw": {"class_name": "9","duration": "25","subject_name": "English Literature"}}}},"hindi_grammar": {"video_lessons": {"हिंदी भाषा": {"-NNHA6axQjjM991cPvW_": {"class_name": "9","duration": "76","subject_name": "हिन्दी व्याकरण"}}}},"history": {"video_lessons": {"The French Revolution": {"-NNHAM0yNF4YcoQSTBRf": {"class_name": "9","duration": "56","subject_name": "History"}}}},"science": {"video_lessons": {"Matter in Our Surroundings": {"-NNHAC_cPxJTZ5iQYj3k": {"class_name": "9","duration": "5","subject_name": "Science"},"-NNHAFNjXkIJjLMuWmQU": {"class_name": "9","duration": "3","subject_name": "Science"}}}}},"hindi": {"computers": {"video_lessons": {"जावास्क्रिप्ट के प्रयोग से क्लाइंट की ओर से स्क्रिप्टिंग": {"-NNHAnYUZNVek70Amlbb": {"class_name": "9","duration": "20","subject_name": "Computers"}}}},"english_literature": {"video_lessons": {"The Fun They Had": {"-NO5EMyVA3PgaxoeYzbQ": {"class_name": "9","duration": "49","subject_name": "English Literature"}}}}}}}}}

现有嵌套循环实现

def reports_data():
    a = input("Enter the UID for the total reports: ")
    # userclass = input("Enter which class data you want to see or write 'all': ")
    total_time = []
    users = db.child("reports").child("app_reports").child(a).child("data").get()
    for i in users.each():
        # user_class = (i.key())
        var_class = (i.val()) #json of the class is being stored in the var variable
        for x, y in var_class.items():
                var_board = (y) #json of the board is being stored in the var variable
                for q, w in var_board.items():
                    var_lang = (w) #json of the language is being stored in the var variable
                    for e, r in var_lang.items():
                        var_subject = (r) #json of the subject is being stored in the var variable
                        for t, s in var_subject.items():
                            var_mode = (s) #json of the subject is being stored in the var variable
                            for u, o in var_mode.items():
                                var_chap = (o) #json of the type is being stored in the var variable
                                for a, p in var_chap.items():
                                    var_mode = (p) #json of the mode is being stored in the var variable
                                    for d, f in var_mode.items():
                                    #  outfile.write(f) #json of the perticular item watched is being stored in the var variable                                                
                                        if d == 'duration':
                                            total_time.append(f)
    total = sum(int(time) for time in total_time)
    print("Total time:", total,"sec")

更高效通用的解决方案

现有方法的问题是硬编码了嵌套层数,一旦JSON结构发生微小变化(比如新增/减少一层嵌套),代码就会失效。推荐使用递归遍历或迭代式遍历的方式,实现通用的字段收集逻辑:

方法1:递归生成器(简洁易读)

通过递归遍历所有嵌套的字典和数组,自动收集所有duration字段的值:

def collect_durations(data):
    # 处理字典类型
    if isinstance(data, dict):
        for key, value in data.items():
            if key == "duration":
                yield int(value)  # 直接转成整数,避免后续重复转换
            else:
                # 递归遍历子节点
                yield from collect_durations(value)
    # 处理数组类型(如果JSON中存在数组结构,兼容扩展)
    elif isinstance(data, list):
        for item in data:
            yield from collect_durations(item)

def reports_data():
    a = input("Enter the UID for the total reports: ")
    users = db.child("reports").child("app_reports").child(a).child("data").get()
    # 将Firebase返回的对象转为字典
    data_dict = users.val()
    # 直接对生成器求和,内存效率更高
    total = sum(collect_durations(data_dict))
    print(f"Total time: {total} sec")

方法2:迭代式遍历(避免递归深度限制)

如果JSON嵌套极深(超过Python默认递归深度限制),可以用栈实现迭代式遍历,避免栈溢出:

def collect_durations_iterative(data):
    stack = [data]
    while stack:
        current = stack.pop()
        if isinstance(current, dict):
            for key, value in current.items():
                if key == "duration":
                    yield int(value)
                else:
                    stack.append(value)
        elif isinstance(current, list):
            for item in current:
                stack.append(item)

def reports_data():
    a = input("Enter the UID for the total reports: ")
    users = db.child("reports").child("app_reports").child(a).child("data").get()
    data_dict = users.val()
    total = sum(collect_durations_iterative(data_dict))
    print(f"Total time: {total} sec")

方案优势

  • 通用性强:无需依赖固定的嵌套层数,自动适配任何深度的JSON结构
  • 代码简洁:相比多层嵌套循环,可读性和可维护性大幅提升
  • 内存高效:使用生成器yield,无需先将所有duration存入列表,直接边遍历边累加,节省内存

内容的提问来源于stack exchange,提问作者Partik Chandra

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 17:12:58