You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python实现特定格式CSV转指定结构JSON的技术问询

问题描述

需要将结构化CSV文件转换为指定格式的JSON文件,要求主位置元素为county(县)名称,子元素为city(城市)名称。

输入CSV示例:

"id","city","diacritice","county","auto","zip","populatie","lat","lng"
"1","Buftea","Buftea","Ilfov","IF","70000","19202","44.5629744","25.9388214"
"2","Buciumeni","Buciumeni","Ilfov","IF","70000","2976","44.5460939","25.9574846"
"3","Otopeni","Otopeni","Ilfov","IF","75100","12540","44.5671874","26.0835113"
"4","Odaile","Odăile","Ilfov","IF","75100","1321","44.5435944","26.0487028"

目标JSON格式:

{
  "locations": [
    {
      "name": "New York",
      "slug": "",
      "description": "",
      "meta": [
        {"image_id" :  45},
        {"_icon" :  ""}
      ],
      "order": 0,
      "child": [
        {"name": "New York", "order": 1 },
        {"name" : "Port Chester", "order": 2},
        {"name" : "Mineola", "order": 3},
        {"name" : "Mount Vernon", "order": 4},
        {"name" : "Hempstead", "order": 5},
        {"name" : "Lynbrook", "order": 6},
        {"name" : "Yonkers", "order": 7},
        {"name" : "Franklin Square", "order": 8}
      ]
    }
  ]
}

用户已编写初步脚本,但无法在JSON结构中嵌套循环获取对应城市,现有脚本如下:

import csv
import json

# First, parse the CSV data using the csv module

county = set()

with open('localitati.csv', 'r', newline='') as csvfile:
  # Use the csv.reader function to read the data line by line
  reader = csv.reader(csvfile)

  # Loop through each line of the CSV data
  for row in reader:
    # Print the name from the second column of the CSV data
    #print(row[3])
    county.add(row[3])

for cou in county:
    csvData = '"1","Buftea","Ilfov","IF","70000","19202","44.5629744","25.9388214"'

    parsedData = list(csv.reader([csvData]))

    # Next, construct the JSON object using the parsed data
    jsonData = {
      "locations": [
        {
          "name": cou,
          "slug": "",
          "description": "",
          "meta": [
            {"image_id" : ""},
            {"_icon" : ""}
          ],
          "order": "",
          "child": [
              #here i tired the for loop to get the cities 
            {"name": parsedData[0][1], "order": ""},
            {"name": parsedData[0][1], "order": ""},
          ]
        }
      ]
    }

# Finally, output the JSON object
print(json.dumps(jsonData, indent=2))

with open("sample.json", "w") as outfile:
    outfile.write(json.dumps(jsonData, indent=2))

完善后的脚本
import csv
import json

def main():
    # 先将CSV数据按县分组存储,键为县名,值为对应城市列表
    county_cities = {}
    
    with open('localitati.csv', 'r', newline='', encoding='utf-8') as csvfile:
        reader = csv.DictReader(csvfile)  # 使用DictReader更易通过列名获取数据
        
        for row in reader:
            county_name = row['county']
            city_name = row['city']
            
            if county_name not in county_cities:
                county_cities[county_name] = []
            county_cities[county_name].append(city_name)
    
    # 构建目标JSON结构
    json_data = {
        "locations": []
    }
    
    # 遍历每个县,生成对应的位置项
    for order, (county_name, cities) in enumerate(county_cities.items(), start=0):
        location = {
            "name": county_name,
            "slug": "",
            "description": "",
            "meta": [
                {"image_id": ""},
                {"_icon": ""}
            ],
            "order": order,
            "child": []
        }
        
        # 为每个城市生成子项,order从1开始递增
        for city_order, city in enumerate(cities, start=1):
            location["child"].append({
                "name": city,
                "order": city_order
            })
        
        json_data["locations"].append(location)
    
    # 输出并保存JSON
    print(json.dumps(json_data, indent=2, ensure_ascii=False))
    with open("sample.json", "w", encoding='utf-8') as outfile:
        outfile.write(json.dumps(json_data, indent=2, ensure_ascii=False))

if __name__ == "__main__":
    main()

关键改进点
  • 使用csv.DictReader替代csv.reader,直接通过列名(如row['county']、row['city'])获取数据,避免索引错误,代码更易读
  • 先一次性读取CSV并按县分组存储城市列表,避免重复读取文件,提升效率
  • 用enumerate自动生成order值,确保顺序正确且无需手动维护
  • 添加encoding='utf-8'参数,支持带重音符号的城市名(如Odăile)正确读写
  • 修复原脚本中仅保留最后一个县数据的问题,所有县都会被加入locations数组

内容的提问来源于stack exchange,提问作者protonusr22

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 06:35:26