Python抓取YouTube下载链接遇href为javascript:void(0)及内容缺失问题求助
基于爬虫的YouTube下载器开发问题汇总
使用站点
- 目标抓取站点:https://www.y2mate.com/
实现思路
利用站点特性,将YouTube视频链接(示例:https://www.youtube.com/watch?v=dQw4w9WgXcQ)拼接在https://www.youtubepp.com/后,跳转后搜索栏会自动填充目标视频,计划用requests请求该拼接链接提取下载地址。
遇到的问题
- 问题1:下载按钮的DOM结构中,480p下载按钮的href属性为
javascript:void(0),完整标签如下:
该标签没有onClick属性,已查询过相关资料仍无法解决如何提取下载链接。<a href="javascript:void(0)" rel="nofollow" type="button" class="btn btn-success" data-toggle="modal" data-target="#progress" data-ftype="mp4" data-fquality="480"> <i class="glyphicon glyphicon-download-alt"></i> Download </a> - 问题2:运行如下代码提取页面HTML后,解析结果中没有需要的下载链接区块:
from bs4 import BeautifulSoup import requests DOMAIN = "https://www.youtubepp.com/" URL = "https://www.youtube.com/watch?v=dQw4w9WgXcQ" download_url = DOMAIN + URL params = { "hl": "en" # 加该参数是因为默认返回非英文内容 } res = requests.get(download_url, params=params) soup = BeautifulSoup(res.text, "html.parser") print(soup.prettify())
替代尝试
因为第一个问题,更换了下载站点,将视频链接中watch?v=dQw4w9WgXcQ部分拼接在https://www.ssyoutube.com/后使用,解决了第一个问题,但仍存在BeautifulSoup解析后无对应内容的问题。
问题排查记录
尝试解决方案后的首次报错
使用测试视频ID运行代码后返回如下错误:
Youtube video id: 9iHM6X6uUH8 Jim Yosef - Link [NCS Release] Traceback (most recent call last): File "c:\Users\rayya\Other\Web_Scraping\Youtube Download\main.py", line 58, in <module> tmpID = re.findall(getId, soup.find("script", {"type": "text/javascript"}).getText())[0] IndexError: list index out of range
错误排查与新报错
排查发现报错行是查找type为text/javascript的script标签的代码,打印soup完整内容确认对应位置存在需要的内容但代码未识别,将代码修改为tmpID = soup.findAll("script", {"type": "text/javascript"})[0]后该问题解决,但出现新报错:
Traceback (most recent call last): File "c:\Users\rayya\Other\Web_Scraping\Youtube Download\main.py", line 91, in <module> toDownload = download.attrs["href"] AttributeError: 'NoneType' object has no attribute 'attrs'
再次打印第二次生成的soup完整内容,暂时无法继续排查。
无正则方案的报错与修复
运行无正则方案返回相同错误,调试信息显示var k__id的值为空:
Jim Yosef - Link [NCS Release] DEBUG: DEBUG: (<- need to bee the value of var k__id) 1: {'quality': '1080p HFR', 'type': 'mp4'} 2: {'quality': '720p HFR', 'type': 'mp4'} 3: {'quality': '480', 'type': 'mp4'} 4: {'quality': '360', 'type': 'mp4'} 5: {'quality': '240p', 'type': 'mp4'} 6: {'quality': '144p', 'type': 'mp4'} 7: {'quality': '144p', 'type': '3gp'} 8: {'quality': '128', 'type': 'mp3'} 9: {'quality': '128', 'type': 'mp3'} Select stream [1-9]: 3 Traceback (most recent call last): File "c:\Users\rayya\Other\Web_Scraping\Youtube Download\sol2.py", line 71, in <module> toDownload = download.attrs["href"] AttributeError: 'NoneType' object has no attribute 'attrs'
判断是.getText()方法返回None导致正则无法匹配结果,尝试自行提取script标签后用正则匹配var k__id的值:
myScript = soup.findAll("script", {"type": "text/javascript"})[0] # 不使用.find().getText()方法 print("DEBUG:", myScript) getId = re.compile(r'(?<=var k__id = ")\w*(?=";)') tmpID = getId.findall(myScript)[0] print("DEBUG:", tmpID, "(<- need to bee the value of var k__id)")
运行后报错:
Traceback (most recent call last): File "c:\Users\rayya\Other\Web_Scraping\Youtube Download\main.py", line 62, in <module> tmpID = getId.findall(myScript)[0] File "C:\Users\rayya\AppData\Local\Programs\Python\Python39\lib\re.py", line 241, in findall return _compile(pattern, flags).findall(string) TypeError: expected string or bytes-like object
打印myScript类型为<class 'bs4.element.Tag'>,将其转为字符串(myScript = str(myScript))后即可正常运行。
内容的提问来源于stack exchange,提问作者Muhd Mairaj
相关产品推荐
相关产品推荐

