You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中通过CSV匹配JSON嵌套地理数据报错求助

问题分析与解决方案

1. 先解决「TypeError: string indices must be integers」错误

这个错误的核心是你遍历的data里,每个x是字符串而非嵌套字典。你需要先确认两点:

  • 必须用json.load()(从文件加载)或json.loads()(从字符串加载)把原始JSON内容转成Python的字典/列表结构,直接读文件得到的是字符串,无法用字典索引。
  • 确认data的结构:如果JSON外层是列表,for x in data遍历的是列表里的字典元素;如果外层是字典,你需要遍历data.values()才能拿到嵌套数据。

正确遍历示例(假设JSON外层是列表,每个元素含size嵌套结构):

import json

# 正确加载JSON数据
with open('your_data.json', 'r') as f:
    data = json.load(f)  # 此时data是列表,每个元素为字典

# 初始化计数器(假设已拿到目标位置的经纬度,注意转成数字)
target_lat = float(lat)
target_long = float(long)
count = 0

for item in data:
    # 先检查嵌套结构是否存在,避免KeyError
    if 'size' in item and 'location' in item['size']:
        current_lat = item['size']['location']['latitude']
        current_long = item['size']['location']['longitude']
        if current_lat == target_lat and current_long == target_long:
            count += 1

2. 优化CSV读取代码

你的CSV读取可以简化,用with语句自动关闭文件更安全,同时把数据转成字典,后续查找位置效率更高:

import csv

# 把CSV转成 {位置名: (纬度, 经度)} 的字典
location_coords = {}
with open('wine.csv', 'r') as f:
    reader = csv.reader(f, delimiter=',')
    next(reader)  # 跳过CSV表头(如果没有表头就删掉这行)
    for row in reader:
        loc_name = row[1]
        lat = float(row[3])
        long = float(row[2])
        location_coords[loc_name] = (lat, long)

3. 处理用户输入(支持最多3个位置+输入验证)

现在实现最多3个位置的输入,并验证输入是否在CSV中存在:

# 获取用户输入,按空格分割后取前3个
user_input = input("请输入最多3个位置,用空格分隔:").strip().split()
target_locations = user_input[:3]

# 筛选有效位置
valid_locations = []
for loc in target_locations:
    if loc in location_coords:
        valid_locations.append(loc)
    else:
        print(f"警告:位置「{loc}」在CSV中不存在,已忽略")

if not valid_locations:
    print("没有有效的位置输入,程序退出")
    exit()

4. 整合完整逻辑(清晰易懂版)

把所有步骤整合,逻辑更清晰,同时避免各种异常:

import json
import csv

def main():
    # 1. 加载CSV位置数据
    location_coords = {}
    with open('wine.csv', 'r') as f:
        reader = csv.reader(f, delimiter=',')
        next(reader)  # 跳过表头
        for row in reader:
            loc_name = row[1]
            lat = float(row[3])
            long = float(row[2])
            location_coords[loc_name] = (lat, long)

    # 2. 处理用户输入与验证
    user_input = input("请输入最多3个位置,用空格分隔:").strip().split()
    target_locations = user_input[:3]
    valid_locations = []
    for loc in target_locations:
        if loc in location_coords:
            valid_locations.append(loc)
        else:
            print(f"警告:位置「{loc}」不存在,已忽略")

    if not valid_locations:
        print("无有效位置,程序结束")
        return

    # 3. 加载JSON数据
    with open('your_data.json', 'r') as f:
        data = json.load(f)

    # 4. 统计每个有效位置的出现次数
    result = {}
    for loc in valid_locations:
        target_lat, target_long = location_coords[loc]
        count = 0
        for item in data:
            # 安全检查嵌套结构的类型和键是否存在
            if (isinstance(item, dict) 
                and 'size' in item 
                and isinstance(item['size'], dict) 
                and 'location' in item['size'] 
                and isinstance(item['size']['location'], dict)):
                current_lat = item['size']['location'].get('latitude')
                current_long = item['size']['location'].get('longitude')
                if current_lat == target_lat and current_long == target_long:
                    count += 1
        # 保存结果
        result[loc] = {
            'Lat': target_lat,
            'Long': target_long,
            'Sales': count
        }

    # 打印最终结果
    print("统计结果:")
    for loc, info in result.items():
        print(f"{loc}: 纬度{info['Lat']}, 经度{info['Long']}, 出现次数{info['Sales']}")

if __name__ == "__main__":
    main()

关键知识点拆解

  • with语句:自动管理文件的打开与关闭,避免忘记关闭文件导致资源泄漏。
  • 字典查找:把CSV数据转成字典后,查找位置的时间从O(n)降到O(1),数据量大时效率提升明显。
  • 类型转换:CSV读取的所有值都是字符串,必须转成float才能和JSON里的数值类型正确比较,否则字符串和数字永远不相等。
  • 安全检查:遍历JSON时先检查每个层级的键和类型,避免出现KeyError或TypeError。
  • 输入验证:过滤无效输入,避免后续统计逻辑出错。

内容的提问来源于stack exchange,提问作者Smashandgrab

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 13:34:38