提取JSON数据entities下的text值报KeyError: 'text'错误如何解决
问题原因
- 你遍历了
entities下所有类型的实体字段,但是Twitter API返回的实体中,不是所有类型的实体都包含text键:比如entities下的media、urls、user_mentions类型实体,对应的文本字段为display_url、expanded_url、screen_name,不存在text字段,直接强制访问就会触发KeyError。 - 非空判断逻辑顺序错误:先执行
for v in value遍历再判断if value,空的value不会走后续逻辑,但遍历动作已经执行,没有起到提前过滤空值的作用。
解决方案
方案1:仅提取包含text字段的实体
使用字典的get方法安全取值,同时调整非空判断顺序,过滤掉无text字段的实体,修正后代码如下:
import json with open('tweets.json') as f: data = json.load(f) for item in data: tweet_id = item.get('id') # 适配你后续入库需要的tweet_id字段 for entity_type, entity_list in item['entities'].items(): # 先判断实体列表非空再遍历 if not entity_list: continue for v in entity_list: s_index = v['indices'][0] e_index = v['indices'][1] # 安全取值,不存在text则返回None all_texts = v.get('text') # 过滤无text的实体,也可按需补充其他类型实体的取值逻辑 if all_texts is None: continue # 后续你的入库逻辑 # cur.execute('INSERT INTO entities (tweet_id, type, value, start_index, end_index) VALUES (?, ?, ?, ?, ?)', (tweet_id, entity_type, all_texts, s_index, e_index))
方案2:适配所有实体类型的文本取值
如果需要提取所有类型实体的对应文本,可以提前配置实体类型和对应文本字段的映射关系,示例如下:
import json # 实体类型和对应文本字段的映射关系,可按需调整 FIELD_MAPPING = { 'hashtags': 'text', 'symbols': 'text', 'user_mentions': 'screen_name', 'urls': 'expanded_url', 'media': 'display_url' } with open('tweets.json') as f: data = json.load(f) for item in data: tweet_id = item.get('id') for entity_type, entity_list in item['entities'].items(): if not entity_list: continue for v in entity_list: s_index = v['indices'][0] e_index = v['indices'][1] # 根据实体类型取对应字段的值 target_field = FIELD_MAPPING.get(entity_type, 'text') all_texts = v.get(target_field) if all_texts is None: continue # 后续入库逻辑
内容的提问来源于stack exchange,提问作者Stainler
相关产品推荐
相关产品推荐

