Python实现特定格式CSV转指定结构JSON的技术问询
问题描述
需要将结构化CSV文件转换为指定格式的JSON文件,要求主位置元素为county(县)名称,子元素为city(城市)名称。
输入CSV示例:
"id","city","diacritice","county","auto","zip","populatie","lat","lng" "1","Buftea","Buftea","Ilfov","IF","70000","19202","44.5629744","25.9388214" "2","Buciumeni","Buciumeni","Ilfov","IF","70000","2976","44.5460939","25.9574846" "3","Otopeni","Otopeni","Ilfov","IF","75100","12540","44.5671874","26.0835113" "4","Odaile","Odăile","Ilfov","IF","75100","1321","44.5435944","26.0487028"
目标JSON格式:
{ "locations": [ { "name": "New York", "slug": "", "description": "", "meta": [ {"image_id" : 45}, {"_icon" : ""} ], "order": 0, "child": [ {"name": "New York", "order": 1 }, {"name" : "Port Chester", "order": 2}, {"name" : "Mineola", "order": 3}, {"name" : "Mount Vernon", "order": 4}, {"name" : "Hempstead", "order": 5}, {"name" : "Lynbrook", "order": 6}, {"name" : "Yonkers", "order": 7}, {"name" : "Franklin Square", "order": 8} ] } ] }
用户已编写初步脚本,但无法在JSON结构中嵌套循环获取对应城市,现有脚本如下:
import csv import json # First, parse the CSV data using the csv module county = set() with open('localitati.csv', 'r', newline='') as csvfile: # Use the csv.reader function to read the data line by line reader = csv.reader(csvfile) # Loop through each line of the CSV data for row in reader: # Print the name from the second column of the CSV data #print(row[3]) county.add(row[3]) for cou in county: csvData = '"1","Buftea","Ilfov","IF","70000","19202","44.5629744","25.9388214"' parsedData = list(csv.reader([csvData])) # Next, construct the JSON object using the parsed data jsonData = { "locations": [ { "name": cou, "slug": "", "description": "", "meta": [ {"image_id" : ""}, {"_icon" : ""} ], "order": "", "child": [ #here i tired the for loop to get the cities {"name": parsedData[0][1], "order": ""}, {"name": parsedData[0][1], "order": ""}, ] } ] } # Finally, output the JSON object print(json.dumps(jsonData, indent=2)) with open("sample.json", "w") as outfile: outfile.write(json.dumps(jsonData, indent=2))
完善后的脚本
import csv import json def main(): # 先将CSV数据按县分组存储,键为县名,值为对应城市列表 county_cities = {} with open('localitati.csv', 'r', newline='', encoding='utf-8') as csvfile: reader = csv.DictReader(csvfile) # 使用DictReader更易通过列名获取数据 for row in reader: county_name = row['county'] city_name = row['city'] if county_name not in county_cities: county_cities[county_name] = [] county_cities[county_name].append(city_name) # 构建目标JSON结构 json_data = { "locations": [] } # 遍历每个县,生成对应的位置项 for order, (county_name, cities) in enumerate(county_cities.items(), start=0): location = { "name": county_name, "slug": "", "description": "", "meta": [ {"image_id": ""}, {"_icon": ""} ], "order": order, "child": [] } # 为每个城市生成子项,order从1开始递增 for city_order, city in enumerate(cities, start=1): location["child"].append({ "name": city, "order": city_order }) json_data["locations"].append(location) # 输出并保存JSON print(json.dumps(json_data, indent=2, ensure_ascii=False)) with open("sample.json", "w", encoding='utf-8') as outfile: outfile.write(json.dumps(json_data, indent=2, ensure_ascii=False)) if __name__ == "__main__": main()
关键改进点
- 使用
csv.DictReader替代csv.reader,直接通过列名(如row['county']、row['city'])获取数据,避免索引错误,代码更易读 - 先一次性读取CSV并按县分组存储城市列表,避免重复读取文件,提升效率
- 用
enumerate自动生成order值,确保顺序正确且无需手动维护 - 添加
encoding='utf-8'参数,支持带重音符号的城市名(如Odăile)正确读写 - 修复原脚本中仅保留最后一个县数据的问题,所有县都会被加入
locations数组
内容的提问来源于stack exchange,提问作者protonusr22
相关产品推荐
相关产品推荐

