非API获取YouTube视频点赞数及爬取链接空集合问题求助
问题
我正在编写程序,目标是获取用户指定内容相关的所有YouTube视频链接,同时不使用API获取这些视频的点赞数。但目前编写的代码运行后仅返回空集合,请求帮忙排查问题。相关代码如下:
import requests from bs4 import BeautifulSoup import tkinter from tkinter import simpledialog root = tkinter.Tk() root.withdraw() content = simpledialog.askstring("Input", "Veuillez saisir une chaîne de caractères :", parent=root) url = f"https://www.youtube.com/results?search_query={content}" response = requests.get(url) soup = BeautifulSoup(response.content, 'html.parser') set_likes = set() set_links = set() #getting number of views for video in soup.find_all("a", href=True): video_href = video['href'] #Verify if the url is correct if "/watch?" in video_href: video_url = f"https://www.youtube.com{video_href}" video_response = requests.get(video_url) video_soup = BeautifulSoup(video_response.content, 'html.parser') #Getting number of likes try: likes = video_soup.find("button", class_="like-button-renderer-like-button").text set_likes.add(likes) set_links.add(video_url) except AttributeError: print(f"Impossible to get links for the videos: {video_url}") # Afficher les résultats print("links:", set_links) print("likes:", set_likes)
问题排查与修复方案
核心问题分析
代码返回空集合主要有以下几个原因:
- 反爬拦截:YouTube会识别非浏览器发起的请求,直接用
requests.get()获取的页面是不含真实内容的静态HTML,无法解析到目标元素。 - 选择器失效:
like-button-renderer-like-button这个类名已被YouTube废弃,当前页面的点赞按钮结构完全不同。 - 无效链接过滤不足:搜索结果页中含
/watch?的<a>标签包含大量导航类无效链接,导致后续请求无法获取有效内容。
修复后的代码
import requests from bs4 import BeautifulSoup import tkinter from tkinter import simpledialog # 模拟浏览器请求头,绕过基础反爬机制 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } root = tkinter.Tk() root.withdraw() content = simpledialog.askstring("Input", "Veuillez saisir une chaîne de caractères :", parent=root) # 对搜索关键词做URL编码,避免特殊字符导致请求错误 encoded_content = requests.utils.quote(content) url = f"https://www.youtube.com/results?search_query={encoded_content}" # 添加请求头获取真实页面内容 response = requests.get(url, headers=headers) soup = BeautifulSoup(response.content, 'html.parser') set_likes = set() set_links = set() # 精准定位搜索结果中的视频链接 video_links = soup.find_all("a", href=True, id="video-title") for video in video_links: video_href = video['href'] if "/watch?" in video_href: video_url = f"https://www.youtube.com{video_href}" # 再次添加请求头获取视频页内容 video_response = requests.get(video_url, headers=headers) video_soup = BeautifulSoup(video_response.content, 'html.parser') # 适配当前YouTube页面结构的点赞数提取逻辑 try: # 通过aria-label属性定位点赞按钮,提取其中的数字 like_btn = video_soup.find("button", {"aria-label": lambda x: x and "j'aime" in x.lower()}) if like_btn: likes = like_btn.get("aria-label").split()[0] set_likes.add(likes) set_links.add(video_url) else: print(f"无法获取视频点赞数: {video_url}") except Exception as e: print(f"处理视频时出错 {video_url}: {str(e)}") # 输出结果 print("links:", set_links) print("likes:", set_likes)
关键修复说明
- 添加请求头:用
User-Agent模拟浏览器请求,避免被YouTube直接拦截。 - URL编码关键词:处理用户输入的特殊字符,确保请求URL合法。
- 精准筛选视频链接:通过
id="video-title"定位真实视频链接,减少无效请求。 - 更新点赞数提取逻辑:利用
aria-label属性定位点赞按钮,提取其中的数字内容,适配当前YouTube页面结构。
内容的提问来源于stack exchange,提问作者Fire35
相关产品推荐
相关产品推荐

