You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从JSON文件中高效提取多类动物数据并生成字典列表

问题描述

需要从JSON文档中提取类型包含cat或snail的动物数据,并导出到单独文件。当前代码仅能提取单类(cat)数据,无法实现多类提取,同时希望找到高效实现方法。

示例数据

{
    "type": "blah",
    "animals": [
        {
            "type": "dog1",
            "name": "oscar"
        },
        {
            "type": "dog2",
            "name": "John"
        },
        {
            "type": "snail1",
            "name": "Fred",
            "Colors": ["Red"],
            "Contact_info": [{"Owner": "Jill", "Owner_number": "123"}]
        },
        {
            "type": "cat3",
            "name": "Freddy",
            "Colors": ["Blue"],
            "Contact_info": [{"Owner": "Ann", "Owner_number": "1323"}]
        }
    ]
}

预期结果

[
    {
        "type": "snail1",
        "name": "Fred",
        "Colors": ["Red"],
        "Contact_info": [{"Owner": "Jill", "Owner_number": "123"}]
    },
    {
        "type": "cat3",
        "name": "Freddy",
        "Colors": ["Blue"],
        "Contact_info": [{"Owner": "Ann", "Owner_number": "1323"}]
    }
]

现有问题代码

import json

with open("data.json", "rb") as f:
    # Load obj list
    data = json.load(f)["animals"]

# Create a list of dictionaries if obj type contains "snail" or cat
animal= [a for a in data if "snail" in a.get("type")]
animal = [a for a in data if "cat" in a.get("type")]

解决方案

1. 修复基础逻辑错误

现有代码的问题在于两次赋值animal变量,第二次的列表推导式完全覆盖了第一次的结果,最终只保留了cat类型的数据。需要把两个条件合并到同一个推导式中:

import json

with open("data.json", "r", encoding="utf-8") as f:
    data = json.load(f)["animals"]

# 合并两个条件,提取包含snail或cat的条目
animals = [a for a in data if "snail" in a.get("type", "") or "cat" in a.get("type", "")]

2. 高效扩展方案(适合多关键词场景)

如果后续需要匹配更多关键词,直接写多个or会冗余。可以用集合存储关键词,通过检查动物类型是否包含集合中的任意关键词实现,扩展性更强,数据量大时效率更高:

import json

# 定义需要匹配的关键词集合
TARGET_TYPES = {"snail", "cat"}

with open("data.json", "r", encoding="utf-8") as f:
    data = json.load(f)["animals"]

# 高效筛选:检查动物type中是否包含任意目标关键词
animals = [
    a for a in data 
    if any(keyword in a.get("type", "") for keyword in TARGET_TYPES)
]

3. 导出结果到文件

提取完成后,用json.dump将结果写入目标文件,设置ensure_ascii=False和indent=4保证中文显示正常且格式美观:

# 导出到结果文件
with open("filtered_animals.json", "w", encoding="utf-8") as f:
    json.dump(animals, f, ensure_ascii=False, indent=4)

内容的提问来源于stack exchange,提问作者user18774110

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 12:18:21