在Google Colab中用Pytube下载YouTube音频遇RegexMatchError求助
解决Colab中pytube的RegexMatchError问题
你遇到的RegexMatchError: __init__: could not find match for ^\w+\W是由于pytube内置的正则表达式无法匹配YouTube当前的页面结构导致的,以下是针对Colab环境的有效解决方法:
解决步骤
1. 更新pytube到最新版本
Colab默认安装的pytube版本可能滞后,先执行升级命令:
!pip install --upgrade pytube
2. 修复正则表达式(若升级后仍报错)
如果升级后问题依旧,说明最新版的正则表达式仍未适配YouTube的页面变化,手动修改pytube的cipher.py文件:
# 替换cipher.py中的目标正则表达式 !sed -i 's/^\\w+\\W/^\\$*\\w+\\W/' /usr/local/lib/python3.10/dist-packages/pytube/cipher.py
注:如果你的pytube安装路径不同,先运行
!pip show pytube查看Location字段,替换上述命令中的路径。
3. 修改代码添加User-Agent(避免被YouTube拦截)
Colab的默认请求头容易被识别为非浏览器请求,添加自定义User-Agent并完善异常处理:
import os from pytube import YouTube # 创建目标文件夹(不存在则自动创建) destination = "/content/mp3_files" os.makedirs(destination, exist_ok=True) # 自定义浏览器User-Agent user_agent = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" for url in url_list: print(url) try: # 初始化YouTube对象时指定User-Agent yt = YouTube( url, use_oauth=False, allow_oauth_cache=False, user_agent=user_agent ) title = yt.title print(f"processing {title}") mp3_path = os.path.join(destination, f"{title}.mp3") if os.path.exists(mp3_path): print("file exists") continue video = yt.streams.filter(only_audio=True).first() print(video) out_file = video.download(output_path=destination) base, ext = os.path.splitext(out_file) new_file = base + '.mp3' os.rename(out_file, new_file) print(f"已成功保存:{new_file}") except Exception as e: print(f"处理{url}时出错:{str(e)}")
内容的提问来源于stack exchange,提问作者PriyankaJ
相关产品推荐
相关产品推荐

