You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用pytube获取小型YouTube频道ID时出现404错误的问询

解决Pytube获取小型YouTube频道ID时的404错误

问题描述

我通过Pytube编写checkID方法获取YouTube频道ID,代码如下:

def checkID(self, channel_url_or_name: str) -> str:
    try:
        print(channel_url_or_name)
        return Channel(channel_url_or_name).channel_id

测试发现,大型频道(如https://www.youtube.com/@Blazedrust)可正常返回channel_id,但小型频道(如https://www.youtube.com/@daveyswerkplaats)会触发404错误,完整报错堆栈:

daveyswerkplaats
https://www.youtube.com/c/daveyswerkplaats
Exception in thread Thread-1 (checkingFeed):
Traceback (most recent call last):
  File "C:\Users\keize\AppData\Local\Programs\Python\Python310\lib\threading.py", line 1016, in _bootstrap_inner  
    self.run()
  File "C:\Users\keize\AppData\Local\Programs\Python\Python310\lib\threading.py", line 953, in run
    self._target(*self._args, **self._kwargs)
  File "E:\projects\Moederjager_bot\cogs\youtube_checker.py", line 63, in checkingFeed
    hiddenId = self.getChannelId(youtube_channel_id)
  File "E:\projects\Moederjager_bot\cogs\youtube_checker.py", line 40, in getChannelId
    return self.checkID(channel_url_or_name)
  File "E:\projects\Moederjager_bot\cogs\youtube_checker.py", line 31, in checkID
    return Channel(channel_url_or_name).channel_id
  File "C:\Users\keize\AppData\Local\Programs\Python\Python310\lib\site-packages\pytube\contrib\channel.py", line 
58, in channel_id
    print(self.initial_data['metadata']['channelMetadataRenderer']['externalId'])
  File "C:\Users\keize\AppData\Local\Programs\Python\Python310\lib\site-packages\pytube\contrib\playlist.py", line 81, in initial_data
    self._initial_data = extract.initial_data(self.html)
  File "C:\Users\keize\AppData\Local\Programs\Python\Python310\lib\site-packages\pytube\contrib\channel.py", line 
79, in html
    self._html = request.get(self.videos_url)
  File "C:\Users\keize\AppData\Local\Programs\Python\Python310\lib\site-packages\pytube\request.py", line 53, in get
    response = _execute_request(url, headers=extra_headers, timeout=timeout)
  File "C:\Users\keize\AppData\Local\Programs\Python\Python310\lib\site-packages\pytube\request.py", line 37, in _execute_request
    return urlopen(request, timeout=timeout)  # nosec
  File "C:\Users\keize\AppData\Local\Programs\Python\Python310\lib\urllib\request.py", line 216, in urlopen       
    return opener.open(url, data, timeout)
  File "C:\Users\keize\AppData\Local\Programs\Python\Python310\lib\urllib\request.py", line 525, in open
    response = meth(req, response)
  File "C:\Users\keize\AppData\Local\Programs\Python\Python310\lib\urllib\request.py", line 634, in http_response 
    response = self.parent.error(
  File "C:\Users\keize\AppData\Local\Programs\Python\Python310\lib\urllib\request.py", line 557, in error
    result = self._call_chain(*args)
  File "C:\Users\keize\AppData\Local\Programs\Python\Python310\lib\urllib\request.py", line 496, in _call_chain   
    result = func(*args)
  File "C:\Users\keize\AppData\Local\Programs\Python\Python310\lib\urllib\request.py", line 749, in http_error_302    return self.parent.open(new, timeout=req.timeout)
  File "C:\Users\keize\AppData\Local\Programs\Python\Python310\lib\urllib\request.py", line 525, in open
    response = meth(req, response)
  File "C:\Users\keize\AppData\Local\Programs\Python\Python310\lib\urllib\request.py", line 634, in http_response 
    response = self.parent.error(
  File "C:\Users\keize\AppData\Local\Programs\Python\Python310\lib\urllib\request.py", line 557, in error
    result = self._call_chain(*args)
  File "C:\Users\keize\AppData\Local\Programs\Python\Python310\lib\urllib\request.py", line 496, in _call_chain   
    result = func(*args)
  File "C:\Users\keize\AppData\Local\Programs\Python\Python310\lib\urllib\request.py", line 749, in http_error_302    return self.parent.open(new, timeout=req.timeout)
  File "C:\Users\keize\AppData\Local\Programs\Python\Python310\lib\urllib\request.py", line 525, in open
    response = meth(req, response)
  File "C:\Users\keize\AppData\Local\Programs\Python\Python310\lib\urllib\request.py", line 634, in http_response 
    response = self.parent.error(
  File "C:\Users\keize\AppData\Local\Programs\Python\Python310\lib\urllib\request.py", line 563, in error
    return self._call_chain(*args)
  File "C:\Users\keize\AppData\Local\Programs\Python\Python310\lib\urllib\request.py", line 496, in _call_chain   
    result = func(*args)
  File "C:\Users\keize\AppData\Local\Programs\Python\Python310\lib\urllib\request.py", line 643, in http_error_default
    raise HTTPError(req.full_url, code, msg, hdrs, fp)
urllib.error.HTTPError: HTTP Error 404: Not Found

目前临时改用了基于requests和BeautifulSoup的实现,但希望修复Pytube原方法的问题。

错误原因

从报错堆栈可知,Pytube的Channel类默认请求频道的视频列表页(self.videos_url)来获取数据,而非频道主页。对于小型频道,YouTube可能因内容过少导致视频列表页不存在,或者路径跳转逻辑与大型频道不同,从而触发404错误。

修复方案

1. 自定义Channel逻辑,优先请求频道主页

绕过Pytube默认的视频列表页请求,直接获取频道主页HTML并提取channel_id:

def checkID(self, channel_url_or_name: str) -> str:
    try:
        from pytube import request
        from pytube.extract import initial_data

        # 统一处理输入格式:支持纯用户名/@格式/完整URL
        if not channel_url_or_name.startswith('http'):
            if channel_url_or_name.startswith('@'):
                channel_url = f"https://www.youtube.com/{channel_url_or_name}"
            else:
                channel_url = f"https://www.youtube.com/@{channel_url_or_name}"
        else:
            channel_url = channel_url_or_name

        # 添加浏览器请求头,避免被YouTube拦截
        headers = {
            'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
        }
        html = request.get(channel_url, headers=headers)
        data = initial_data(html)
        return data['metadata']['channelMetadataRenderer']['externalId']
    except Exception as e:
        print(f"获取频道ID失败: {str(e)}")
        # 失败时回退到你的requests+BeautifulSoup方案
        return self.get_channel_id_from_url(channel_url)

2. 升级Pytube到最新版本

旧版本Pytube对YouTube新频道格式(如@用户名)的适配存在bug,执行以下命令升级:

pip install --upgrade pytube

新版本可能已修复小型频道的404问题。

3. 修改Pytube源码(临时方案)

若不想自定义方法,可直接修改Pytube的Channel类实现:
找到pytube/contrib/channel.py文件,将html属性的请求目标从视频列表页改为频道主页:

@property
def html(self):
    if not self._html:
        # 替换原请求videos_url的代码
        self._html = request.get(self.channel_url)
    return self._html

注意:源码修改会在Pytube升级后被覆盖,仅作临时使用。

替代方案(优化版)

给你现有的requests+BeautifulSoup方法添加请求头,提升稳定性:

import requests
from bs4 import BeautifulSoup
import re

def get_channel_id_from_url(url):
    headers = {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
    }
    response = requests.get(url, headers=headers)
    if response.status_code == 200:
        soup = BeautifulSoup(response.content, 'html.parser')
        meta_tag = soup.find('meta', itemprop='channelId')
        if meta_tag:
            return meta_tag['content']
        link_tags = soup.find_all('link', rel='canonical')
        for link in link_tags:
            if 'channel' in link['href']:
                channel_id_match = re.search(r'\/channel\/([^\/]+)', link['href'])
                if channel_id_match:
                    return channel_id_match.group(1)
    return None

内容的提问来源于stack exchange,提问作者ki5os

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 22:55:54