You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用urllib获取超过20条YouTube频道搜索结果?Python脚本求助

解决YouTube关键词搜索获取超过20个频道URL的问题

方案一:使用YouTube Data API(推荐)

这是官方支持的稳定方案,不会因页面结构变动失效,还能轻松实现分页获取大量结果。

步骤:

  1. 前往Google Cloud平台创建项目,启用YouTube Data API v3,并获取API密钥。
  2. 使用google-api-python-client库发送搜索请求,通过pageToken参数实现分页,单次最多请求50条结果。

代码示例:

from googleapiclient.discovery import build

# 替换为你的API密钥
API_KEY = "你的API密钥"
YOUTUBE_API_SERVICE_NAME = "youtube"
YOUTUBE_API_VERSION = "v3"

def search_youtube_channels(keyword, max_results=100):
    youtube = build(YOUTUBE_API_SERVICE_NAME, YOUTUBE_API_VERSION, developerKey=API_KEY)
    
    channel_urls = []
    next_page_token = None
    
    while len(channel_urls) < max_results:
        # 控制单次请求数量,不超过剩余需求和API上限50
        current_limit = min(max_results - len(channel_urls), 50)
        
        search_response = youtube.search().list(
            q=keyword,
            part="id",
            type="channel",
            maxResults=current_limit,
            pageToken=next_page_token
        ).execute()
        
        # 提取频道ID并生成完整URL
        for item in search_response["items"]:
            channel_id = item["id"]["channelId"]
            channel_urls.append(f"https://www.youtube.com/channel/{channel_id}")
        
        # 获取下一页标识,无则停止
        next_page_token = search_response.get("nextPageToken")
        if not next_page_token:
            break
    
    return channel_urls

if __name__ == "__main__":
    search_keyword = input("Search Keyword \n")
    results = search_youtube_channels(search_keyword, 100)
    for url in results:
        print(url)

注意事项:

  • 安装依赖:执行pip install google-api-python-client
  • API有调用配额限制,具体配额可在Google Cloud控制台查看。

方案二:模拟动态翻页爬取(不推荐)

YouTube搜索结果为动态加载,默认仅返回第一页(约20条)。若不用API,需模拟浏览器滚动加载或构造AJAX请求,但这种方式易触发反爬,且页面结构变动会导致代码失效。

使用Selenium模拟浏览器(示例):

Selenium可以模拟用户滚动页面加载更多内容,适合快速验证,但不适合长期使用。

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import time

def search_channels_with_selenium(keyword, max_results=100):
    # 需提前安装对应浏览器的驱动,这里以Chrome为例
    driver = webdriver.Chrome()
    search_url = f"https://www.youtube.com/results?search_query={keyword}&sp=EgIQAg%3D%3D"
    driver.get(search_url)
    
    channel_urls = set()  # 用集合去重
    
    while len(channel_urls) < max_results:
        # 等待频道链接加载完成
        channel_elements = WebDriverWait(driver, 10).until(
            EC.presence_of_all_elements_located((By.CSS_SELECTOR, "a.channel-link"))
        )
        
        # 提取所有频道URL
        for elem in channel_elements:
            href = elem.get_attribute("href")
            if href and "/channel/" in href:
                channel_urls.add(href)
        
        # 已达到目标数量则停止
        if len(channel_urls) >= max_results:
            break
        
        # 滚动到页面底部加载更多内容
        driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
        time.sleep(3)  # 等待加载完成
        
        # 检查是否还有更多结果
        try:
            no_more_elem = driver.find_element(By.CSS_SELECTOR, "ytd-search-pyv-renderer #message")
            if "No more results" in no_more_elem.text:
                break
        except:
            pass
    
    driver.quit()
    # 返回指定数量的结果
    return list(channel_urls)[:max_results]

if __name__ == "__main__":
    search_keyword = input("Search Keyword \n")
    results = search_channels_with_selenium(search_keyword, 100)
    for url in results:
        print(url)

注意事项:

  • 安装依赖:执行pip install selenium,并下载对应浏览器的驱动(如ChromeDriver)。
  • 频繁请求可能触发YouTube反爬机制,建议添加随机延迟、使用代理或无头浏览器。

内容的提问来源于stack exchange,提问作者Prashant Ranjan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 09:20:42