You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将大量JSON数据批量导入Django模型对应字段?

Django批量导入JSON数据到模型字段的解决方案

问题背景

现有包含大量数据的JSON文件,已创建对应结构的Django模型,需要将JSON中每个数据项的字段映射到模型的对应字段(而非使用JSONField),手动导入耗时过长,需批量处理方案。

JSON数据示例

[
    {
        "end_year": "",
        "intensity": 6,
        "sector": "Energy",
        "topic": "gas",
        "insight": "Annual Energy Outlook",
        "url": "http://example.com",
        "region": "Northern America",
        "start_year": "",
        "impact": "",
        "added": "January, 20 2017 03:51:25",
        "published": "January, 09 2017 00:00:00",
        "country": "United States of America",
        "relevance": 2,
        "pestle": "Industries",
        "source": "EIA",
        "title": "Test phase",
        "likelihood": 3
    }
]

Django模型代码

from django.db import models

class myModel(models.Model):
    end_year = models.PositiveIntegerField(null=True, blank=True)
    intensity = models.PositiveSmallIntegerField(null=True, blank=True)
    sector = models.CharField(max_length=255, null=True, blank=True)
    topic = models.CharField(max_length=55, null=True, blank=True)
    insight = models.TextField()
    url = models.URLField(max_length=300)
    region = models.CharField(max_length=50, null=True, blank=True)
    start_year = models.PositiveIntegerField(null=True, blank=True)
    impact = models.CharField(max_length=255, null=True, blank=True)
    added = models.DateTimeField(null=True, blank=True)
    published = models.DateTimeField(null=True, blank=True)
    country = models.CharField(max_length=50, null=True, blank=True)
    relevance = models.PositiveIntegerField(null=True, blank=True)
    pestle = models.CharField(max_length=100, null=True, blank=True)
    source = models.CharField(max_length=200, null=True, blank=True)
    title = models.CharField(max_length=300, null=True, blank=True)
    likelihood = models.PositiveIntegerField(null=True, blank=True)

批量导入方案

方案1:使用Django Shell快速处理

这是最直接的临时处理方式,步骤如下:

  1. 打开Django Shell:
python manage.py shell
  1. 执行以下导入脚本(替换JSON文件路径和你的app名称):
import json
from datetime import datetime
from yourapp.models import myModel  # 替换为你的app名称

# 读取JSON文件
with open('path/to/your/data.json', 'r', encoding='utf-8') as f:
    data = json.load(f)

# 处理数据并准备批量创建的对象列表
objects_to_create = []
for item in data:
    # 转换空字符串为None,适配模型的null=True字段
    processed_item = {k: None if v == "" else v for k, v in item.items()}
    
    # 处理日期字段:将字符串转为datetime对象
    if processed_item['added']:
        processed_item['added'] = datetime.strptime(processed_item['added'], "%B, %d %Y %H:%M:%S")
    if processed_item['published']:
        processed_item['published'] = datetime.strptime(processed_item['published'], "%B, %d %Y %H:%M:%S")
    
    # 转换数字字段(空值已处理为None,避免转换报错)
    num_fields = ['end_year', 'start_year', 'intensity', 'relevance', 'likelihood']
    for field in num_fields:
        if processed_item[field] is not None:
            processed_item[field] = int(processed_item[field])
    
    # 创建模型对象(不立即保存)
    objects_to_create.append(myModel(**processed_item))

# 批量保存,提高效率
myModel.objects.bulk_create(objects_to_create, batch_size=1000)

方案2:自定义Django管理命令(适合重复执行)

如果需要多次导入或在生产环境使用,建议编写自定义管理命令:

  1. 在你的app目录下创建management/commands目录,结构如下:
yourapp/
├── management/
│   ├── __init__.py
│   └── commands/
│       ├── __init__.py
│       └── import_json_data.py
  1. 在import_json_data.py中写入以下代码(替换你的app名称):
import json
from datetime import datetime
from django.core.management.base import BaseCommand
from yourapp.models import myModel  # 替换为你的app名称

class Command(BaseCommand):
    help = '批量导入JSON数据到myModel'

    def add_arguments(self, parser):
        parser.add_argument('json_file', type=str, help='JSON文件的路径')

    def handle(self, *args, **options):
        json_path = options['json_file']
        with open(json_path, 'r', encoding='utf-8') as f:
            data = json.load(f)
        
        objects_to_create = []
        for item in data:
            processed_item = {k: None if v == "" else v for k, v in item.items()}
            
            # 日期转换
            if processed_item['added']:
                processed_item['added'] = datetime.strptime(processed_item['added'], "%B, %d %Y %H:%M:%S")
            if processed_item['published']:
                processed_item['published'] = datetime.strptime(processed_item['published'], "%B, %d %Y %H:%M:%S")
            
            # 数字转换
            num_fields = ['end_year', 'start_year', 'intensity', 'relevance', 'likelihood']
            for field in num_fields:
                if processed_item[field] is not None:
                    processed_item[field] = int(processed_item[field])
            
            objects_to_create.append(myModel(**processed_item))
        
        myModel.objects.bulk_create(objects_to_create, batch_size=1000)
        self.stdout.write(self.style.SUCCESS(f'成功导入{len(objects_to_create)}条数据'))
  1. 执行命令导入数据:
python manage.py import_json_data path/to/your/data.json

注意事项

  • 数据验证:导入前先抽样检查JSON数据,确保字段类型和模型匹配,避免因格式错误导致导入失败。
  • 批量大小:batch_size可根据服务器内存调整,过大可能导致内存占用过高,过小则降低导入效率。
  • 事务处理:若需保证数据一致性,可使用transaction.atomic()包裹批量创建逻辑。

内容的提问来源于stack exchange,提问作者Dhiman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.30 19:23:15