JSON过滤报错TypeError: string indices must be integers 求助
解决Elasticsearch返回JSON的过滤问题
问题原因
你遇到的TypeError: string indices must be integers是因为直接遍历了Elasticsearch返回的顶层字典,拿到的是字典的键(字符串类型),而非实际的文档数据。Elasticsearch查询返回有固定层级结构,实际文档数据嵌套在hits.hits数组中,每个文档的业务数据又存放在_source字段下。
解决方案
典型Elasticsearch返回结构参考
你的JSON文件结构应该类似以下常规ES返回格式:
{ "took": 2, "timed_out": false, "_shards": { "total": 1, "successful": 1, "skipped": 0, "failed": 0 }, "hits": { "total": { "value": 5, "relation": "eq" }, "max_score": 1.0, "hits": [ { "_index": "your_index", "_id": "1", "_source": { "cbaCodeParts": { "HHH": "300", "otherField": "xxx" }, ... } }, ... ] } }
代码修正
场景1:保留完整ES响应结构(包含_index、_score等元数据)
import json with open('2022-10.json', 'r') as f: es_response = json.load(f) # 定位到实际的文档数组 doc_list = es_response['hits']['hits'] # 过滤cbaCodeParts.HHH不等于'300'的文档 filtered_docs = [ doc for doc in doc_list if doc['_source']['cbaCodeParts']['HHH'] != '300' ] # 将过滤后的文档放回原响应结构 filtered_response = es_response.copy() filtered_response['hits']['hits'] = filtered_docs # 输出格式化后的结果 print(json.dumps(filtered_response, indent=2))
场景2:仅提取业务数据(_source部分)
如果不需要ES的元数据,只保留业务字段:
import json with open('2022-10.json', 'r') as f: es_response = json.load(f) doc_list = es_response['hits']['hits'] filtered_sources = [ doc['_source'] for doc in doc_list if doc['_source']['cbaCodeParts']['HHH'] != '300' ] print(json.dumps(filtered_sources, indent=2))
额外注意事项
- 如果你的JSON结构和典型结构有差异,先打印查看实际层级:
import pprint with open('2022-10.json', 'r') as f: es_response = json.load(f) pprint.pprint(es_response) - 若存在部分文档缺失
cbaCodeParts或HHH字段的情况,为避免KeyError,可添加判断:filtered_docs = [ doc for doc in doc_list if 'cbaCodeParts' in doc['_source'] and doc['_source']['cbaCodeParts'].get('HHH') != '300' ]
内容的提问来源于stack exchange,提问作者Stanislav Jirak
相关产品推荐
相关产品推荐

