如何用urllib获取超过20条YouTube频道搜索结果?Python脚本求助
解决YouTube关键词搜索获取超过20个频道URL的问题
方案一:使用YouTube Data API(推荐)
这是官方支持的稳定方案,不会因页面结构变动失效,还能轻松实现分页获取大量结果。
步骤:
- 前往Google Cloud平台创建项目,启用YouTube Data API v3,并获取API密钥。
- 使用
google-api-python-client库发送搜索请求,通过pageToken参数实现分页,单次最多请求50条结果。
代码示例:
from googleapiclient.discovery import build # 替换为你的API密钥 API_KEY = "你的API密钥" YOUTUBE_API_SERVICE_NAME = "youtube" YOUTUBE_API_VERSION = "v3" def search_youtube_channels(keyword, max_results=100): youtube = build(YOUTUBE_API_SERVICE_NAME, YOUTUBE_API_VERSION, developerKey=API_KEY) channel_urls = [] next_page_token = None while len(channel_urls) < max_results: # 控制单次请求数量,不超过剩余需求和API上限50 current_limit = min(max_results - len(channel_urls), 50) search_response = youtube.search().list( q=keyword, part="id", type="channel", maxResults=current_limit, pageToken=next_page_token ).execute() # 提取频道ID并生成完整URL for item in search_response["items"]: channel_id = item["id"]["channelId"] channel_urls.append(f"https://www.youtube.com/channel/{channel_id}") # 获取下一页标识,无则停止 next_page_token = search_response.get("nextPageToken") if not next_page_token: break return channel_urls if __name__ == "__main__": search_keyword = input("Search Keyword \n") results = search_youtube_channels(search_keyword, 100) for url in results: print(url)
注意事项:
- 安装依赖:执行
pip install google-api-python-client - API有调用配额限制,具体配额可在Google Cloud控制台查看。
方案二:模拟动态翻页爬取(不推荐)
YouTube搜索结果为动态加载,默认仅返回第一页(约20条)。若不用API,需模拟浏览器滚动加载或构造AJAX请求,但这种方式易触发反爬,且页面结构变动会导致代码失效。
使用Selenium模拟浏览器(示例):
Selenium可以模拟用户滚动页面加载更多内容,适合快速验证,但不适合长期使用。
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import time def search_channels_with_selenium(keyword, max_results=100): # 需提前安装对应浏览器的驱动,这里以Chrome为例 driver = webdriver.Chrome() search_url = f"https://www.youtube.com/results?search_query={keyword}&sp=EgIQAg%3D%3D" driver.get(search_url) channel_urls = set() # 用集合去重 while len(channel_urls) < max_results: # 等待频道链接加载完成 channel_elements = WebDriverWait(driver, 10).until( EC.presence_of_all_elements_located((By.CSS_SELECTOR, "a.channel-link")) ) # 提取所有频道URL for elem in channel_elements: href = elem.get_attribute("href") if href and "/channel/" in href: channel_urls.add(href) # 已达到目标数量则停止 if len(channel_urls) >= max_results: break # 滚动到页面底部加载更多内容 driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") time.sleep(3) # 等待加载完成 # 检查是否还有更多结果 try: no_more_elem = driver.find_element(By.CSS_SELECTOR, "ytd-search-pyv-renderer #message") if "No more results" in no_more_elem.text: break except: pass driver.quit() # 返回指定数量的结果 return list(channel_urls)[:max_results] if __name__ == "__main__": search_keyword = input("Search Keyword \n") results = search_channels_with_selenium(search_keyword, 100) for url in results: print(url)
注意事项:
- 安装依赖:执行
pip install selenium,并下载对应浏览器的驱动(如ChromeDriver)。 - 频繁请求可能触发YouTube反爬机制,建议添加随机延迟、使用代理或无头浏览器。
内容的提问来源于stack exchange,提问作者Prashant Ranjan
相关产品推荐
相关产品推荐

