使用IMDB API批量调用时列表推导出现KeyError(year)问题
问题
我正在开发一个展示和朋友多年来共同观看电影评分的项目,使用RapidAPI上API Dojo的IMDB API。流程是先导入指定列的Excel表格,定义FilmData函数通过IMDB ID获取电影元数据、导演等信息,再结合两人的评分追加数据到表格中。现在用三个列表(IDS、listamby、listmattis)通过列表推导批量调用该函数时,第5条数据后出现KeyError: year,但手动在RapidAPI中调用对应ID一切正常,怀疑是调用频率限制导致的问题,求解答。
代码如下:
import pandas as pd import requests import json kolonner = ['ID' , 'ÅR' , 'TITEL' , 'RATING_AMBY' , 'RATING_MATTIS' , 'SAMLET_RATING' , 'STUDIE' , 'INSTRUKTØR' , 'GENRE' , 'RUNTIME'] df = pd.read_excel(r'C:\Users\Downloads\FilmlisteNY.xlsx' , header = 0 , index_col=0 , sheet_name='Sheet1') #Definerer funktion def FilmData(film , RATING_AMBY , RATING_MATTIS): #HENTER FØRSTE RUNDE OPLYSNINGER BASERET PÅ ID url = "https://imdb8.p.rapidapi.com/title/get-meta-data" querystring = {"ids":film} headers = { "X-RapidAPI-Key": "Can't show you key, but it is here", "X-RapidAPI-Host": "imdb8.p.rapidapi.com" } response = requests.request("GET" , url , headers=headers , params=querystring) response = response.json() # Henter informationer omkring DIRECTOR url2 = "https://imdb8.p.rapidapi.com/title/get-full-credits" querystring2 = {"tconst":film} response2 = requests.request("GET" , url2 , headers=headers , params=querystring2) response2 = response2.json() #GEMMER INFORMATIONER I GLOBALE VARIABLE global ID ID = film global ÅR ÅR = response[film]['title']['year'] #if response[film]['title']['year'] in response else None global TITEL TITEL = response[film]['title']['title'] #if response[film]['title']['title'] in response else None global GENRE GENRE = response[film]['genres'][0] #if response[film]['genres'][0] in response else None global RUNTIME RUNTIME = response[film]['title']['runningTimeInMinutes'] #if response[film]['title']['runningTimeInMinutes'] in response else None global INSTRUKTØR INSTRUKTØR = response2['crew']['director'][0]['name'] #if response2['crew']['director'][0]['name'] in response else None global RATING_SAMLET RATING_SAMLET = (RATING_AMBY + RATING_MATTIS) / 2 list = [ID , ÅR , TITEL , RATING_AMBY , RATING_MATTIS , RATING_SAMLET , 0 , INSTRUKTØR , GENRE , RUNTIME] global df df = df.append(pd.DataFrame([list] , columns = kolonner) , ignore_index = True) listamby = [4, 5, 4, 3, 4, 5, 3, 3, 3, 2, 1, 3, 3, 4, 4, 3, 2, 4, 1, 5, 5, 3, 5, 5, 4, 3, 4, 5, 1, 1, 4, 4, 3, 4, 4, 1, 5, 5, 2, 2, 5, 5, 4, 1, 4, 4, 3, 3, 2, 4, 4, 4, 3, 2, 3, 2, 4, 1, 2, 2, 3, 4, 1, 1, 3, 4, 1, 3, 1, 3] listmattis = [3, 5, 4, 3, 4, 5, 2, 3, 4, 2, 1, 2, 2, 4, 2, 2, 3, 4, 1, 5, 5, 3, 4, 5, 4, 4, 4, 3, 3, 1, 4, 4, 4, 5, 4, 3, 5, 4, 2, 1, 5, 5, 3, 1, 4, 5, 4, 4, 2, 4, 3, 4, 3, 2, 4, 2, 4, 1, 3, 2, 4, 4, 2, 1, 3, 4, 2, 3, 1, 3] IDS = ["tt2105044","tt0407887","tt7784604","tt0012299","tt0070224","tt10126662","tt1259521","tt0123755","tt2321549","tt0758742","tt5059406", "tt4733640", "tt4163224", "tt0078767", "tt0112735", "tt3654796", "tt4399952", "tt0386846", "tt7043012", "tt1375666", "tt0117381", "tt0079643", "tt1677733", "tt0211130", "tt0443706", "tt0004832", "tt0104800", "tt0098374", "tt0395584", "tt0440803", "tt0464141", "tt3958034", "tt0014199", "tt0816692", "tt0101674", "tt1703199", "tt0059043", "tt0215632", "tt1320244", "tt0805570", "tt0182313", "tt0327056", "tt0065387", "tt1662293", "tt1442053", "tt0309698", "tt0333952", "tt0335106", "tt21336716", "tt7349950", "tt2386278", "tt0430359", "tt0103181", "tt0069345", "tt1119191", "tt2450186", "tt0057664", "tt0147630", "tt1285009", "tt7153766", "tt2388715", "tt0324114", "tt0816556", "tt0007574", "tt0361862", "tt5430018", "tt0424404", "tt0486822", "tt0370848", "tt6114050"] [FilmData(x,y,z) for x,y,z in zip(IDS,listamby,listmattis)]
解决方案
1. 先验证调用频率限制问题
批量调用时触发API限流的概率很高,限流后API可能返回空数据或错误响应,导致读取year字段时触发KeyError。可以做这两步:
- 检查响应状态码:在请求后添加判断,确认请求是否成功:
response = requests.request("GET", url, headers=headers, params=querystring) if response.status_code != 200: print(f"ID {film} 请求失败,状态码: {response.status_code}") return response = response.json() - 添加请求延迟:导入
time模块,在两次调用之间加入延迟,降低调用频率:import time # 在函数末尾添加 time.sleep(1)
2. 修复字段读取的安全问题
即使没有限流,部分电影数据可能缺失字段,直接链式取值容易报错。改用字典get方法做安全取值,不存在的字段返回None:
# 替换原来的全局变量赋值代码 ÅR = response.get(film, {}).get('title', {}).get('year', None) TITEL = response.get(film, {}).get('title', {}).get('title', None) GENRE = response.get(film, {}).get('genres', [None])[0] RUNTIME = response.get(film, {}).get('title', {}).get('runningTimeInMinutes', None) INSTRUKTØR = response2.get('crew', {}).get('director', [{}])[0].get('name', None)
3. 优化代码逻辑
- 去掉
global变量:让函数返回数据行,再统一追加到DataFrame,避免全局变量的副作用:def FilmData(film, RATING_AMBY, RATING_MATTIS): # ... 原有请求逻辑 ... RATING_SAMLET = (RATING_AMBY + RATING_MATTIS) / 2 return [film, ÅR, TITEL, RATING_AMBY, RATING_MATTIS, RATING_SAMLET, 0, INSTRUKTØR, GENRE, RUNTIME] - 改用循环+异常捕获:代替列表推导,方便定位错误:
for x,y,z in zip(IDS,listamby,listmattis): try: row = FilmData(x,y,z) df = df.append(pd.DataFrame([row], columns=kolonner), ignore_index=True) except Exception as e: print(f"处理ID {x} 时出错: {str(e)}")
内容的提问来源于stack exchange,提问作者Amby95
相关产品推荐
相关产品推荐

