You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Pytube Channel爬取YouTube频道时video_urls失效报IndexError

Pytube Channel类video_urls返回空列表触发IndexError的解决方法

问题描述

使用Pytube的Channel类爬取YouTube频道时,video_urls返回空列表,执行c.video_urls[:20]切片操作时触发IndexError,其他功能正常。

代码示例

from pytube import Channel, YouTube
import re

# 假设slicer和reversed_slicer为自定义处理函数
def slicer(text, start):
    return text.split(start)[1] if start in text else ""

def reversed_slicer(text, end):
    return text.rsplit(end)[0] if end in text else ""

c = Channel("https://www.youtube.com/c/BuildEmpire")
print(f"This is a test number of video urls:  {len(c.video_urls)}")
for url in c.video_urls[:20]:
  video = YouTube(url)
  description2 = video.description
  try:
    videos_credits = slicer(description2,"Video Credits:")
    videos_credits = reversed_slicer(videos_credits,"Thumbnail:")
    urls = re.findall(r'(https?://[^\s]+)', videos_credits)
  except Exception:
    print("Sorry but I choose bad youtube video I will try it again")

报错信息

PS C:\Users\Lukas\Dokumenty\python_scripts\Billionare livestyle> & "c:/Users/Lukas/Dokumenty/python_scripts/Billionare livestyle/env/youtube/Scripts/python.exe" "c:/Users/Lukas/Dokumenty/python_scripts/Billionare livestyle/bilionare_life.py"
[]
0
Traceback (most recent call last):
  File "C:\Users\Lukas\Dokumenty\python_scripts\Billionare livestyle\env\youtube\lib\site-packages\pytube\helpers.py", line 57, in __getitem__
    next_item = next(self.gen)
StopIteration

During handling of the above exception, another exception occurred:

Traceback (most recent call last):
  File "c:\Users\Lukas\Dokumenty\python_scripts\Billionare livestyle\bilionare_life.py", line 198, in <module>
    for url in c.video_urls[:20]:
  File "C:\Users\Lukas\Dokumenty\python_scripts\Billionare livestyle\env\youtube\lib\site-packages\pytube\helpers.py", line 60, in __getitem__
    raise IndexError
IndexError

原因分析

  1. 版本兼容问题:YouTube频繁更新页面结构,旧版本Pytube的解析逻辑失效,无法正确提取视频URL。
  2. URL格式问题:/c/格式的频道URL存在解析bug,YouTube当前更推荐@username或/channel/[ID]格式的URL。
  3. 反爬限制:YouTube检测到非浏览器请求,限制数据返回,导致无法获取视频列表。
  4. 生成器特性:video_urls是延迟加载的生成器,无数据时执行切片会直接触发IndexError。

解决方案

1. 更新Pytube到最新版本

执行命令升级Pytube,修复已知解析bug:

pip install --upgrade pytube

2. 更换频道URL格式

将/c/格式替换为@username格式,Pytube对该格式解析更稳定:

c = Channel("https://www.youtube.com/@BuildEmpire")  # 替换原URL

3. 提前处理空列表,避免切片报错

将video_urls转为列表后判断长度,再执行切片:

video_urls = list(c.video_urls)
print(f"This is a test number of video urls:  {len(video_urls)}")

if not video_urls:
    print("未获取到视频URL,请检查频道URL或Pytube版本")
else:
    for url in video_urls[:20]:
        # 后续处理代码

4. 改用Channel的videos属性直接获取视频对象

videos属性返回视频对象列表,无需重新初始化YouTube实例,同时规避生成器空值问题:

c = Channel("https://www.youtube.com/@BuildEmpire")
videos = list(c.videos)
print(f"This is a test number of videos:  {len(videos)}")

for video in videos[:20]:
    try:
        description2 = video.description
        videos_credits = slicer(description2,"Video Credits:")
        videos_credits = reversed_slicer(videos_credits,"Thumbnail:")
        urls = re.findall(r'(https?://[^\s]+)', videos_credits)
    except Exception as e:
        print(f"处理视频 {video.watch_url} 出错: {str(e)}")

5. 添加自定义请求头绕过反爬

设置浏览器请求头,模拟正常访问:

from pytube.request import set_header

# 模拟Chrome浏览器请求头
set_header('User-Agent', 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36')

c = Channel("https://www.youtube.com/@BuildEmpire")
# 后续代码不变

内容的提问来源于stack exchange,提问作者ALex Break

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 15:40:37