You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Google Colab中用Pytube下载YouTube音频遇RegexMatchError求助

解决Colab中pytube的RegexMatchError问题

你遇到的RegexMatchError: __init__: could not find match for ^\w+\W是由于pytube内置的正则表达式无法匹配YouTube当前的页面结构导致的,以下是针对Colab环境的有效解决方法:

解决步骤

1. 更新pytube到最新版本

Colab默认安装的pytube版本可能滞后,先执行升级命令:

!pip install --upgrade pytube

2. 修复正则表达式(若升级后仍报错)

如果升级后问题依旧,说明最新版的正则表达式仍未适配YouTube的页面变化,手动修改pytube的cipher.py文件:

# 替换cipher.py中的目标正则表达式
!sed -i 's/^\\w+\\W/^\\$*\\w+\\W/' /usr/local/lib/python3.10/dist-packages/pytube/cipher.py

注:如果你的pytube安装路径不同,先运行!pip show pytube查看Location字段,替换上述命令中的路径。

3. 修改代码添加User-Agent(避免被YouTube拦截)

Colab的默认请求头容易被识别为非浏览器请求,添加自定义User-Agent并完善异常处理:

import os
from pytube import YouTube

# 创建目标文件夹(不存在则自动创建)
destination = "/content/mp3_files"
os.makedirs(destination, exist_ok=True)

# 自定义浏览器User-Agent
user_agent = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"

for url in url_list:
    print(url)
    try:
        # 初始化YouTube对象时指定User-Agent
        yt = YouTube(
            url,
            use_oauth=False,
            allow_oauth_cache=False,
            user_agent=user_agent
        )
        title = yt.title
        print(f"processing {title}")
        
        mp3_path = os.path.join(destination, f"{title}.mp3")
        if os.path.exists(mp3_path):
            print("file exists")
            continue  
        
        video = yt.streams.filter(only_audio=True).first()
        print(video)    
        out_file = video.download(output_path=destination)
        
        base, ext = os.path.splitext(out_file)
        new_file = base + '.mp3'
        os.rename(out_file, new_file)
        print(f"已成功保存:{new_file}")
    except Exception as e:
        print(f"处理{url}时出错:{str(e)}")

内容的提问来源于stack exchange,提问作者PriyankaJ

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 19:37:25