Python中通过CSV匹配JSON嵌套地理数据报错求助
问题分析与解决方案
1. 先解决「TypeError: string indices must be integers」错误
这个错误的核心是你遍历的data里,每个x是字符串而非嵌套字典。你需要先确认两点:
- 必须用
json.load()(从文件加载)或json.loads()(从字符串加载)把原始JSON内容转成Python的字典/列表结构,直接读文件得到的是字符串,无法用字典索引。 - 确认
data的结构:如果JSON外层是列表,for x in data遍历的是列表里的字典元素;如果外层是字典,你需要遍历data.values()才能拿到嵌套数据。
正确遍历示例(假设JSON外层是列表,每个元素含size嵌套结构):
import json # 正确加载JSON数据 with open('your_data.json', 'r') as f: data = json.load(f) # 此时data是列表,每个元素为字典 # 初始化计数器(假设已拿到目标位置的经纬度,注意转成数字) target_lat = float(lat) target_long = float(long) count = 0 for item in data: # 先检查嵌套结构是否存在,避免KeyError if 'size' in item and 'location' in item['size']: current_lat = item['size']['location']['latitude'] current_long = item['size']['location']['longitude'] if current_lat == target_lat and current_long == target_long: count += 1
2. 优化CSV读取代码
你的CSV读取可以简化,用with语句自动关闭文件更安全,同时把数据转成字典,后续查找位置效率更高:
import csv # 把CSV转成 {位置名: (纬度, 经度)} 的字典 location_coords = {} with open('wine.csv', 'r') as f: reader = csv.reader(f, delimiter=',') next(reader) # 跳过CSV表头(如果没有表头就删掉这行) for row in reader: loc_name = row[1] lat = float(row[3]) long = float(row[2]) location_coords[loc_name] = (lat, long)
3. 处理用户输入(支持最多3个位置+输入验证)
现在实现最多3个位置的输入,并验证输入是否在CSV中存在:
# 获取用户输入,按空格分割后取前3个 user_input = input("请输入最多3个位置,用空格分隔:").strip().split() target_locations = user_input[:3] # 筛选有效位置 valid_locations = [] for loc in target_locations: if loc in location_coords: valid_locations.append(loc) else: print(f"警告:位置「{loc}」在CSV中不存在,已忽略") if not valid_locations: print("没有有效的位置输入,程序退出") exit()
4. 整合完整逻辑(清晰易懂版)
把所有步骤整合,逻辑更清晰,同时避免各种异常:
import json import csv def main(): # 1. 加载CSV位置数据 location_coords = {} with open('wine.csv', 'r') as f: reader = csv.reader(f, delimiter=',') next(reader) # 跳过表头 for row in reader: loc_name = row[1] lat = float(row[3]) long = float(row[2]) location_coords[loc_name] = (lat, long) # 2. 处理用户输入与验证 user_input = input("请输入最多3个位置,用空格分隔:").strip().split() target_locations = user_input[:3] valid_locations = [] for loc in target_locations: if loc in location_coords: valid_locations.append(loc) else: print(f"警告:位置「{loc}」不存在,已忽略") if not valid_locations: print("无有效位置,程序结束") return # 3. 加载JSON数据 with open('your_data.json', 'r') as f: data = json.load(f) # 4. 统计每个有效位置的出现次数 result = {} for loc in valid_locations: target_lat, target_long = location_coords[loc] count = 0 for item in data: # 安全检查嵌套结构的类型和键是否存在 if (isinstance(item, dict) and 'size' in item and isinstance(item['size'], dict) and 'location' in item['size'] and isinstance(item['size']['location'], dict)): current_lat = item['size']['location'].get('latitude') current_long = item['size']['location'].get('longitude') if current_lat == target_lat and current_long == target_long: count += 1 # 保存结果 result[loc] = { 'Lat': target_lat, 'Long': target_long, 'Sales': count } # 打印最终结果 print("统计结果:") for loc, info in result.items(): print(f"{loc}: 纬度{info['Lat']}, 经度{info['Long']}, 出现次数{info['Sales']}") if __name__ == "__main__": main()
关键知识点拆解
with语句:自动管理文件的打开与关闭,避免忘记关闭文件导致资源泄漏。- 字典查找:把CSV数据转成字典后,查找位置的时间从O(n)降到O(1),数据量大时效率提升明显。
- 类型转换:CSV读取的所有值都是字符串,必须转成
float才能和JSON里的数值类型正确比较,否则字符串和数字永远不相等。 - 安全检查:遍历JSON时先检查每个层级的键和类型,避免出现
KeyError或TypeError。 - 输入验证:过滤无效输入,避免后续统计逻辑出错。
内容的提问来源于stack exchange,提问作者Smashandgrab
相关产品推荐
相关产品推荐

