Django使用Pandas加载JSON数据至模型失败求助
Django加载JSON数据到Jsondata模型的错误修复与替代方案
问题分析
导致加载失败的核心问题:
- 模型方法拼写错误:
__str_应为__str__,缺少末尾下划线会导致对象显示异常。 - 管理命令方法缩进错误:
add_arguments和handle方法未缩进在Command类内部,Django无法识别这些命令逻辑。 - 数据类型不匹配:JSON中
start_year是空字符串,但模型定义为IntegerField,无法直接存储非整数值。 - JSON格式无效:示例JSON条目末尾多了逗号,且未包裹在数组中,会触发
json.load解析失败。
修复后的代码
1. 修正后的Jsondata模型
from django.db import models class Jsondata(models.Model): region = models.CharField(max_length=50, blank=True) # 若允许年份为空,添加null=True;否则需确保JSON中start_year为有效整数 start_year = models.IntegerField(null=True, blank=True) published = models.CharField(max_length=100, blank=True) country = models.CharField(max_length=100, blank=True) def __str__(self): # 修复拼写错误 return self.country
2. 修正后的load_data.py管理命令
import json import pandas as pd from django.core.management.base import BaseCommand from users.models import Jsondata class Command(BaseCommand): help = 'Load data from JSON file using pandas' # 修正缩进:属于Command类的内部方法 def add_arguments(self, parser): parser.add_argument('json_file', type=str) # 修正缩进:属于Command类的内部方法 def handle(self, *args, **kwargs): json_file_path = kwargs['json_file'] with open(json_file_path, 'r') as json_file: # 确保JSON是数组格式:[{"region":...}, {...}] data = json.load(json_file) df = pd.DataFrame(data) for index, row in df.iterrows(): # 处理start_year的空值/非整数情况 start_year = row['start_year'] if start_year == '' or pd.isna(start_year): start_year = None # 对应模型的null=True设置 else: try: start_year = int(start_year) except ValueError: start_year = None # 无法转成整数时设为null item = Jsondata( region=row['region'], start_year=start_year, published=row['published'], country=row['country'] ) item.save() self.stdout.write(self.style.SUCCESS('Data loaded successfully'))
3. 修正后的JSON文件格式
必须是数组包裹的有效JSON,无多余逗号:
[ { "region": "Central America", "start_year": "", "published": "January, 18 2017 00:00:00", "country": "Mexico" }, { "region": "North America", "start_year": "2020", "published": "March, 05 2020 12:00:00", "country": "United States" } ]
替代方案
方案1:用bulk_create提高批量插入效率
如果数据量较大,直接用Django的bulk_create比逐行save性能更优:
import json from django.core.management.base import BaseCommand from users.models import Jsondata class Command(BaseCommand): help = 'Load data from JSON file directly' def add_arguments(self, parser): parser.add_argument('json_file', type=str) def handle(self, *args, **kwargs): json_file_path = kwargs['json_file'] with open(json_file_path, 'r') as json_file: data = json.load(json_file) items = [] for entry in data: start_year = entry['start_year'] # 处理空值和类型转换 if not start_year: start_year = None else: try: start_year = int(start_year) except ValueError: start_year = None items.append(Jsondata( region=entry['region'], start_year=start_year, published=entry['published'], country=entry['country'] )) Jsondata.objects.bulk_create(items) self.stdout.write(self.style.SUCCESS(f'Successfully loaded {len(items)} records'))
方案2:使用Django内置loaddata命令
将JSON转换为Django fixture格式,无需自定义命令即可加载:
[ { "model": "users.jsondata", "pk": 1, "fields": { "region": "Central America", "start_year": null, "published": "January, 18 2017 00:00:00", "country": "Mexico" } } ]
将文件保存到users/fixtures/jsondata.json,执行命令:
python manage.py loaddata jsondata.json
内容的提问来源于stack exchange,提问作者Shreyash mishra
相关产品推荐
相关产品推荐

