使用pytube获取小型YouTube频道ID时出现404错误的问询
解决Pytube获取小型YouTube频道ID时的404错误
问题描述
我通过Pytube编写checkID方法获取YouTube频道ID,代码如下:
def checkID(self, channel_url_or_name: str) -> str: try: print(channel_url_or_name) return Channel(channel_url_or_name).channel_id
测试发现,大型频道(如https://www.youtube.com/@Blazedrust)可正常返回channel_id,但小型频道(如https://www.youtube.com/@daveyswerkplaats)会触发404错误,完整报错堆栈:
daveyswerkplaats https://www.youtube.com/c/daveyswerkplaats Exception in thread Thread-1 (checkingFeed): Traceback (most recent call last): File "C:\Users\keize\AppData\Local\Programs\Python\Python310\lib\threading.py", line 1016, in _bootstrap_inner self.run() File "C:\Users\keize\AppData\Local\Programs\Python\Python310\lib\threading.py", line 953, in run self._target(*self._args, **self._kwargs) File "E:\projects\Moederjager_bot\cogs\youtube_checker.py", line 63, in checkingFeed hiddenId = self.getChannelId(youtube_channel_id) File "E:\projects\Moederjager_bot\cogs\youtube_checker.py", line 40, in getChannelId return self.checkID(channel_url_or_name) File "E:\projects\Moederjager_bot\cogs\youtube_checker.py", line 31, in checkID return Channel(channel_url_or_name).channel_id File "C:\Users\keize\AppData\Local\Programs\Python\Python310\lib\site-packages\pytube\contrib\channel.py", line 58, in channel_id print(self.initial_data['metadata']['channelMetadataRenderer']['externalId']) File "C:\Users\keize\AppData\Local\Programs\Python\Python310\lib\site-packages\pytube\contrib\playlist.py", line 81, in initial_data self._initial_data = extract.initial_data(self.html) File "C:\Users\keize\AppData\Local\Programs\Python\Python310\lib\site-packages\pytube\contrib\channel.py", line 79, in html self._html = request.get(self.videos_url) File "C:\Users\keize\AppData\Local\Programs\Python\Python310\lib\site-packages\pytube\request.py", line 53, in get response = _execute_request(url, headers=extra_headers, timeout=timeout) File "C:\Users\keize\AppData\Local\Programs\Python\Python310\lib\site-packages\pytube\request.py", line 37, in _execute_request return urlopen(request, timeout=timeout) # nosec File "C:\Users\keize\AppData\Local\Programs\Python\Python310\lib\urllib\request.py", line 216, in urlopen return opener.open(url, data, timeout) File "C:\Users\keize\AppData\Local\Programs\Python\Python310\lib\urllib\request.py", line 525, in open response = meth(req, response) File "C:\Users\keize\AppData\Local\Programs\Python\Python310\lib\urllib\request.py", line 634, in http_response response = self.parent.error( File "C:\Users\keize\AppData\Local\Programs\Python\Python310\lib\urllib\request.py", line 557, in error result = self._call_chain(*args) File "C:\Users\keize\AppData\Local\Programs\Python\Python310\lib\urllib\request.py", line 496, in _call_chain result = func(*args) File "C:\Users\keize\AppData\Local\Programs\Python\Python310\lib\urllib\request.py", line 749, in http_error_302 return self.parent.open(new, timeout=req.timeout) File "C:\Users\keize\AppData\Local\Programs\Python\Python310\lib\urllib\request.py", line 525, in open response = meth(req, response) File "C:\Users\keize\AppData\Local\Programs\Python\Python310\lib\urllib\request.py", line 634, in http_response response = self.parent.error( File "C:\Users\keize\AppData\Local\Programs\Python\Python310\lib\urllib\request.py", line 557, in error result = self._call_chain(*args) File "C:\Users\keize\AppData\Local\Programs\Python\Python310\lib\urllib\request.py", line 496, in _call_chain result = func(*args) File "C:\Users\keize\AppData\Local\Programs\Python\Python310\lib\urllib\request.py", line 749, in http_error_302 return self.parent.open(new, timeout=req.timeout) File "C:\Users\keize\AppData\Local\Programs\Python\Python310\lib\urllib\request.py", line 525, in open response = meth(req, response) File "C:\Users\keize\AppData\Local\Programs\Python\Python310\lib\urllib\request.py", line 634, in http_response response = self.parent.error( File "C:\Users\keize\AppData\Local\Programs\Python\Python310\lib\urllib\request.py", line 563, in error return self._call_chain(*args) File "C:\Users\keize\AppData\Local\Programs\Python\Python310\lib\urllib\request.py", line 496, in _call_chain result = func(*args) File "C:\Users\keize\AppData\Local\Programs\Python\Python310\lib\urllib\request.py", line 643, in http_error_default raise HTTPError(req.full_url, code, msg, hdrs, fp) urllib.error.HTTPError: HTTP Error 404: Not Found
目前临时改用了基于requests和BeautifulSoup的实现,但希望修复Pytube原方法的问题。
错误原因
从报错堆栈可知,Pytube的Channel类默认请求频道的视频列表页(self.videos_url)来获取数据,而非频道主页。对于小型频道,YouTube可能因内容过少导致视频列表页不存在,或者路径跳转逻辑与大型频道不同,从而触发404错误。
修复方案
1. 自定义Channel逻辑,优先请求频道主页
绕过Pytube默认的视频列表页请求,直接获取频道主页HTML并提取channel_id:
def checkID(self, channel_url_or_name: str) -> str: try: from pytube import request from pytube.extract import initial_data # 统一处理输入格式:支持纯用户名/@格式/完整URL if not channel_url_or_name.startswith('http'): if channel_url_or_name.startswith('@'): channel_url = f"https://www.youtube.com/{channel_url_or_name}" else: channel_url = f"https://www.youtube.com/@{channel_url_or_name}" else: channel_url = channel_url_or_name # 添加浏览器请求头,避免被YouTube拦截 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } html = request.get(channel_url, headers=headers) data = initial_data(html) return data['metadata']['channelMetadataRenderer']['externalId'] except Exception as e: print(f"获取频道ID失败: {str(e)}") # 失败时回退到你的requests+BeautifulSoup方案 return self.get_channel_id_from_url(channel_url)
2. 升级Pytube到最新版本
旧版本Pytube对YouTube新频道格式(如@用户名)的适配存在bug,执行以下命令升级:
pip install --upgrade pytube
新版本可能已修复小型频道的404问题。
3. 修改Pytube源码(临时方案)
若不想自定义方法,可直接修改Pytube的Channel类实现:
找到pytube/contrib/channel.py文件,将html属性的请求目标从视频列表页改为频道主页:
@property def html(self): if not self._html: # 替换原请求videos_url的代码 self._html = request.get(self.channel_url) return self._html
注意:源码修改会在Pytube升级后被覆盖,仅作临时使用。
替代方案(优化版)
给你现有的requests+BeautifulSoup方法添加请求头,提升稳定性:
import requests from bs4 import BeautifulSoup import re def get_channel_id_from_url(url): headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } response = requests.get(url, headers=headers) if response.status_code == 200: soup = BeautifulSoup(response.content, 'html.parser') meta_tag = soup.find('meta', itemprop='channelId') if meta_tag: return meta_tag['content'] link_tags = soup.find_all('link', rel='canonical') for link in link_tags: if 'channel' in link['href']: channel_id_match = re.search(r'\/channel\/([^\/]+)', link['href']) if channel_id_match: return channel_id_match.group(1) return None
内容的提问来源于stack exchange,提问作者ki5os
相关产品推荐
相关产品推荐

