You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用IMDB API批量调用时列表推导出现KeyError(year)问题

问题

我正在开发一个展示和朋友多年来共同观看电影评分的项目,使用RapidAPI上API Dojo的IMDB API。流程是先导入指定列的Excel表格,定义FilmData函数通过IMDB ID获取电影元数据、导演等信息,再结合两人的评分追加数据到表格中。现在用三个列表(IDS、listamby、listmattis)通过列表推导批量调用该函数时,第5条数据后出现KeyError: year,但手动在RapidAPI中调用对应ID一切正常,怀疑是调用频率限制导致的问题,求解答。

代码如下:

import pandas as pd
import requests
import json

kolonner = ['ID' , 'ÅR' , 'TITEL' , 'RATING_AMBY' , 'RATING_MATTIS' , 'SAMLET_RATING' , 'STUDIE' , 'INSTRUKTØR' , 'GENRE' , 'RUNTIME']
df = pd.read_excel(r'C:\Users\Downloads\FilmlisteNY.xlsx' , header = 0 , index_col=0 , sheet_name='Sheet1')

#Definerer funktion
def FilmData(film , RATING_AMBY , RATING_MATTIS):
    #HENTER FØRSTE RUNDE OPLYSNINGER BASERET PÅ ID
    url = "https://imdb8.p.rapidapi.com/title/get-meta-data"

    querystring = {"ids":film}

    headers = {
        "X-RapidAPI-Key": "Can't show you key, but it is here",
        "X-RapidAPI-Host": "imdb8.p.rapidapi.com"
    }

    response = requests.request("GET" , url , headers=headers , params=querystring)
    response = response.json()


    # Henter informationer omkring DIRECTOR

    url2 = "https://imdb8.p.rapidapi.com/title/get-full-credits"

    querystring2 = {"tconst":film}

    response2 = requests.request("GET" , url2 , headers=headers , params=querystring2)
    response2 = response2.json()

    #GEMMER INFORMATIONER I GLOBALE VARIABLE
    global ID 
    ID = film
    global ÅR
    ÅR = response[film]['title']['year'] #if response[film]['title']['year'] in response else None 
    global TITEL
    TITEL = response[film]['title']['title'] #if response[film]['title']['title'] in response else None 
    global GENRE
    GENRE = response[film]['genres'][0] #if response[film]['genres'][0] in response else None 
    global RUNTIME
    RUNTIME = response[film]['title']['runningTimeInMinutes'] #if response[film]['title']['runningTimeInMinutes'] in response else None 
    global INSTRUKTØR
    INSTRUKTØR = response2['crew']['director'][0]['name'] #if response2['crew']['director'][0]['name'] in response else None 
    global RATING_SAMLET
    RATING_SAMLET = (RATING_AMBY + RATING_MATTIS) / 2
    list = [ID , ÅR , TITEL , RATING_AMBY , RATING_MATTIS , RATING_SAMLET , 0 , INSTRUKTØR , GENRE , RUNTIME]
    global df
    df = df.append(pd.DataFrame([list] , columns = kolonner) , ignore_index = True)

listamby = [4,  5,  4,  3,  4,  5,  3,  3,  3,  2,  1,  3,  3,  4,  4,  3,  2,  4,  1,  5,  5,  3,  5,  5,  4,  3,  4,  5,  1,  1,  4,  4,  3,  4,  4,  1,  5,  5,  2,  2,  5,  5,  4,  1,  4,  4,  3,  3,  2,  4,  4,  4,  3,  2,  3,  2,  4,  1,  2,  2,  3,  4,  1,  1,  3,  4,  1,  3,  1,  3]
listmattis = [3,    5,  4,  3,  4,  5,  2,  3,  4,  2,  1,  2,  2,  4,  2,  2,  3,  4,  1,  5,  5,  3,  4,  5,  4,  4,  4,  3,  3,  1,  4,  4,  4,  5,  4,  3,  5,  4,  2,  1,  5,  5,  3,  1,  4,  5,  4,  4,  2,  4,  3,  4,  3,  2,  4,  2,  4,  1,  3,  2,  4,  4,  2,  1,  3,  4,  2,  3,  1,  3]
IDS = ["tt2105044","tt0407887","tt7784604","tt0012299","tt0070224","tt10126662","tt1259521","tt0123755","tt2321549","tt0758742","tt5059406",    "tt4733640",    "tt4163224",    "tt0078767",    "tt0112735",    "tt3654796",    "tt4399952",    "tt0386846",    "tt7043012",    "tt1375666",    "tt0117381",    "tt0079643",    "tt1677733",    "tt0211130",    "tt0443706",    "tt0004832",    "tt0104800",    "tt0098374",    "tt0395584",    "tt0440803",    "tt0464141",    "tt3958034",    "tt0014199",    "tt0816692",    "tt0101674",    "tt1703199",    "tt0059043",    "tt0215632",    "tt1320244",    "tt0805570",    "tt0182313",    "tt0327056",    "tt0065387",    "tt1662293",    "tt1442053",    "tt0309698",    "tt0333952",    "tt0335106",    "tt21336716",   "tt7349950",    "tt2386278",    "tt0430359",    "tt0103181",    "tt0069345",    "tt1119191",    "tt2450186",    "tt0057664",    "tt0147630",    "tt1285009",    "tt7153766",    "tt2388715",    "tt0324114",    "tt0816556",    "tt0007574",    "tt0361862",    "tt5430018",    "tt0424404",    "tt0486822",    "tt0370848",    "tt6114050"]

[FilmData(x,y,z) for x,y,z in zip(IDS,listamby,listmattis)]
解决方案

1. 先验证调用频率限制问题

批量调用时触发API限流的概率很高,限流后API可能返回空数据或错误响应,导致读取year字段时触发KeyError。可以做这两步:

  • 检查响应状态码:在请求后添加判断,确认请求是否成功:
    response = requests.request("GET", url, headers=headers, params=querystring)
    if response.status_code != 200:
        print(f"ID {film} 请求失败,状态码: {response.status_code}")
        return
    response = response.json()
    
  • 添加请求延迟:导入time模块,在两次调用之间加入延迟,降低调用频率:
    import time
    # 在函数末尾添加
    time.sleep(1)
    

2. 修复字段读取的安全问题

即使没有限流,部分电影数据可能缺失字段,直接链式取值容易报错。改用字典get方法做安全取值,不存在的字段返回None:

# 替换原来的全局变量赋值代码
ÅR = response.get(film, {}).get('title', {}).get('year', None)
TITEL = response.get(film, {}).get('title', {}).get('title', None)
GENRE = response.get(film, {}).get('genres', [None])[0]
RUNTIME = response.get(film, {}).get('title', {}).get('runningTimeInMinutes', None)
INSTRUKTØR = response2.get('crew', {}).get('director', [{}])[0].get('name', None)

3. 优化代码逻辑

  • 去掉global变量:让函数返回数据行,再统一追加到DataFrame,避免全局变量的副作用:
    def FilmData(film, RATING_AMBY, RATING_MATTIS):
        # ... 原有请求逻辑 ...
        RATING_SAMLET = (RATING_AMBY + RATING_MATTIS) / 2
        return [film, ÅR, TITEL, RATING_AMBY, RATING_MATTIS, RATING_SAMLET, 0, INSTRUKTØR, GENRE, RUNTIME]
    
  • 改用循环+异常捕获:代替列表推导,方便定位错误:
    for x,y,z in zip(IDS,listamby,listmattis):
        try:
            row = FilmData(x,y,z)
            df = df.append(pd.DataFrame([row], columns=kolonner), ignore_index=True)
        except Exception as e:
            print(f"处理ID {x} 时出错: {str(e)}")
    

内容的提问来源于stack exchange,提问作者Amby95

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 19:50:25