调用YouTube API获取频道字幕时触发KeyError及后续TypeError问题求助
问题:YouTube字幕统计程序运行异常解决
几个月前开发的程序,用于获取指定YouTube频道所有视频的字幕并统计特定词汇(如firework)的出现次数,当时运行完全正常,如今重新使用时无法正常工作。
原代码
from apiclient.discovery import build from youtube_transcript_api import YouTubeTranscriptApi from pytube import YouTube import os # Your Google API Key api_key = "Your Youtube API here" # The Youtube channel you want to extract data from channel_id = "Youtube channel here" # Youtube Build youtube = build('youtube', 'v3', developerKey=api_key) def get_channel_videos(channel_id): # get Uploads playlist id res = youtube.channels().list(id=channel_id, part='contentDetails').execute() playlist_id = res['items'][0]['contentDetails']['relatedPlaylists']['uploads'] videos = [] next_page_token = None while 1: res = youtube.playlistItems().list(playlistId=playlist_id, part='snippet', maxResults=50, pageToken=next_page_token).execute() videos += res['items'] next_page_token = res.get('nextPageToken') if next_page_token is None: break return videos videos = get_channel_videos(channel_id) video_ids = [] # list of all video_id of channel video_titles = [] # Word to search for word1 = "firework" # Counter for the text files counter = 1 for video in videos: video_ids.append(video['snippet']['resourceId']['videoId']) # Loops through every video on the youtube channel for video_id in video_ids: try: # Gets the transcripts of youtube videos. Only gets English transcripts. responses = YouTubeTranscriptApi.get_transcript( video_id, languages=['en']) # Get's the youtube video's link yt = YouTube("https://www.youtube.com/watch?v="+str(video_id)) #print('\n'+"Video: "+"https://www.youtube.com/watch?v="+str(video_id)+'\n'+'\n'+"Captions:") print("\n"+"Video Number: "+str(counter)+'\n'+"Video: "+"https://www.youtube.com/watch?v="+str(video_id)+'\n'+'Title: '+yt.title+'\n'+"Firework/s Wordcount:") # This is for text file generation filename = "video #" + str(counter) + ".txt" # Generates text file with open(filename, "w") as file: # Writes the title of the Youtube video at the top of the text file file.write(yt.title) # Whitespace file.write("\n") # Writes all the data grabbed from the transcripts into a text file # in text form. for i in responses: file.write("{}\n".format(i)) counter += 1 file.close() # This reads data from all the generated files, # then converts all the data within the file to lowercase, # then it counts the amount of times that the word you are searching # for is used within each file that is generated. fileTest = open(filename, "r") read_data = fileTest.read() word_count = read_data.lower().count(word1) file.close() # Prints out the word count. print(word_count) for response in responses: text = response['text'] except Exception as e: print(e)
首次运行错误(KeyError)
Traceback (most recent call last): File "D:\Downloads\YouTube-Caption-Project-main\main.py", line 47, in <module> videos = get_channel_videos(channel_id) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "D:\Downloads\YouTube-Caption-Project-main\main.py", line 30, in get_channel_videos playlist_id = res['items'][0]['contentDetails']['relatedPlaylists']['uploads'] ~~~^^^^^^^^^ KeyError: 'items'
参考相关问题修改后,出现新错误(TypeError):
Traceback (most recent call last): File "D:\Downloads\YouTube-Caption-Project-main\main.py", line 62, in <module> for video in videos: TypeError: 'NoneType' object is not iterable
解决方案
1. 排查API调用失败根源
- 验证API密钥有效性:登录Google Cloud控制台,确认YouTube Data API v3已启用,API密钥未过期、配额未耗尽,且没有IP访问限制。
- 检查频道ID正确性:确保
channel_id是YouTube频道的原始ID(可通过频道页面源码或第三方工具获取),而非自定义URL或用户名。 - 增加API错误处理:在调用API后先判断响应是否包含有效数据,避免直接索引不存在的字段。
2. 修复get_channel_videos函数
修改函数,确保无论成功或失败都返回可迭代的列表,避免NoneType报错:
def get_channel_videos(channel_id): try: # 获取频道上传列表ID res = youtube.channels().list(id=channel_id, part='contentDetails').execute() if not res.get('items'): print("无法获取频道信息,请检查频道ID或API密钥") return [] playlist_id = res['items'][0]['contentDetails']['relatedPlaylists']['uploads'] videos = [] next_page_token = None while True: res = youtube.playlistItems().list( playlistId=playlist_id, part='snippet', maxResults=50, pageToken=next_page_token ).execute() if res.get('items'): videos += res['items'] next_page_token = res.get('nextPageToken') if next_page_token is None: break return videos except Exception as e: print(f"获取视频列表时出错: {str(e)}") return []
3. 优化字幕统计逻辑(可选)
无需生成文本文件再读取统计,直接在获取字幕时计算,提升效率:
# 替换原统计部分代码 full_text = ' '.join([item['text'].lower() for item in responses]) word_count = full_text.count(word1.lower()) print(word_count)
同时,with语句会自动关闭文件,无需手动调用file.close(),可删除代码中多余的关闭操作。
4. 更新依赖库
YouTube相关库可能因平台结构更新失效,执行以下命令更新:
pip install --upgrade pytube youtube-transcript-api google-api-python-client
5. 细化异常捕获
针对字幕相关异常单独处理,提升错误提示清晰度:
from youtube_transcript_api import NoTranscriptFound, TranscriptsDisabled # 在循环中替换原try-except块 try: responses = YouTubeTranscriptApi.get_transcript(video_id, languages=['en']) # ... 其他代码 ... except NoTranscriptFound: print(f"视频 {video_id} 无英文字幕") except TranscriptsDisabled: print(f"视频 {video_id} 字幕已禁用") except Exception as e: print(f"处理视频 {video_id} 时出错: {str(e)}")
内容的提问来源于stack exchange,提问作者Keenonthedaywalker
相关产品推荐
相关产品推荐

