如何使用pytube获取YouTube个人库中所有播放列表的URL?
获取YouTube个人库中播放列表URL的实现方案
说明
pytube 本身不支持直接访问需要登录权限的YouTube个人库内容,因为个人库的播放列表属于用户私有/需授权的资源。要实现需求,需要借助模拟浏览器登录或者YouTube Data API来完成。
方案一:使用Selenium模拟登录并爬取
这种方式通过模拟浏览器登录YouTube,直接解析个人库页面的播放列表链接:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import time # 初始化Chrome浏览器(需提前下载对应版本的ChromeDriver) driver = webdriver.Chrome() driver.get("https://www.youtube.com/feed/library") # 手动完成YouTube登录(可自行添加自动填充账号密码逻辑,注意账号安全) input("请在浏览器中完成登录,登录后按回车继续...") # 等待播放列表区域加载完成 wait = WebDriverWait(driver, 10) playlist_elements = wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, "ytd-playlist-thumbnail a#thumbnail"))) # 提取所有播放列表URL playlists = [elem.get_attribute("href") for elem in playlist_elements] # 输出结果 for i, playlist_url in enumerate(playlists, start=1): print(f'playlist {i}: {playlist_url}') # 关闭浏览器 driver.quit()
方案二:使用YouTube Data API(更规范的方式)
通过官方API获取用户的播放列表,需提前完成以下准备:
- 前往Google Cloud Console创建项目,启用YouTube Data API v3
- 创建OAuth 2.0客户端ID,下载凭据文件(命名为
credentials.json) - 安装依赖库:
pip install google-api-python-client google-auth-oauthlib
代码示例:
from googleapiclient.discovery import build from google_auth_oauthlib.flow import InstalledAppFlow from google.auth.transport.requests import Request import pickle import os # 定义API权限范围 SCOPES = ["https://www.googleapis.com/auth/youtube.readonly"] def get_authenticated_service(): creds = None # 读取已保存的登录凭据 if os.path.exists("token.pickle"): with open("token.pickle", "rb") as token: creds = pickle.load(token) # 无有效凭据则触发登录流程 if not creds or not creds.valid: if creds and creds.expired and creds.refresh_token: creds.refresh(Request()) else: flow = InstalledAppFlow.from_client_secrets_file( "credentials.json", SCOPES ) creds = flow.run_local_server(port=0) # 保存凭据供下次使用 with open("token.pickle", "wb") as token: pickle.dump(creds, token) return build("youtube", "v3", credentials=creds) def get_user_playlists(youtube): playlists = [] next_page_token = None # 分页获取所有播放列表 while True: request = youtube.playlists().list( part="snippet,contentDetails", mine=True, # 获取当前登录用户的播放列表 maxResults=50, pageToken=next_page_token ) response = request.execute() # 拼接播放列表URL for item in response["items"]: playlist_url = f"https://www.youtube.com/playlist?list={item['id']}" playlists.append(playlist_url) next_page_token = response.get("nextPageToken") if not next_page_token: break return playlists if __name__ == "__main__": youtube = get_authenticated_service() playlists = get_user_playlists(youtube) for i, playlist_url in enumerate(playlists, start=1): print(f'playlist {i}: {playlist_url}')
注意事项
- 使用Selenium时,需保证浏览器驱动与浏览器版本匹配,注意登录状态的维护
- 使用YouTube Data API时,需遵守Google的API配额限制,避免超出调用次数
内容的提问来源于stack exchange,提问作者Rozhyar Gaylan
相关产品推荐
相关产品推荐

