You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

非API获取YouTube视频点赞数及爬取链接空集合问题求助

问题

我正在编写程序,目标是获取用户指定内容相关的所有YouTube视频链接,同时不使用API获取这些视频的点赞数。但目前编写的代码运行后仅返回空集合,请求帮忙排查问题。相关代码如下:

import requests
from bs4 import BeautifulSoup
import tkinter
from tkinter import simpledialog

root = tkinter.Tk()
root.withdraw()


content = simpledialog.askstring("Input", "Veuillez saisir une chaîne de caractères :", parent=root)

url = f"https://www.youtube.com/results?search_query={content}"

response = requests.get(url)
soup = BeautifulSoup(response.content, 'html.parser')

set_likes = set()
set_links = set()

#getting number of views
for video in soup.find_all("a", href=True):
    video_href = video['href']
    #Verify if the url is correct
    if "/watch?" in video_href:
        video_url = f"https://www.youtube.com{video_href}"
        video_response = requests.get(video_url)
        video_soup = BeautifulSoup(video_response.content, 'html.parser')
        
        #Getting number of likes
        try:
            likes = video_soup.find("button", class_="like-button-renderer-like-button").text
            set_likes.add(likes)
            set_links.add(video_url)
        except AttributeError:
            print(f"Impossible to get links for the videos: {video_url}")

# Afficher les résultats
print("links:", set_links)
print("likes:", set_likes)
问题排查与修复方案

核心问题分析

代码返回空集合主要有以下几个原因:

  • 反爬拦截:YouTube会识别非浏览器发起的请求,直接用requests.get()获取的页面是不含真实内容的静态HTML,无法解析到目标元素。
  • 选择器失效:like-button-renderer-like-button这个类名已被YouTube废弃,当前页面的点赞按钮结构完全不同。
  • 无效链接过滤不足:搜索结果页中含/watch?的<a>标签包含大量导航类无效链接,导致后续请求无法获取有效内容。

修复后的代码

import requests
from bs4 import BeautifulSoup
import tkinter
from tkinter import simpledialog

# 模拟浏览器请求头,绕过基础反爬机制
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
}

root = tkinter.Tk()
root.withdraw()

content = simpledialog.askstring("Input", "Veuillez saisir une chaîne de caractères :", parent=root)
# 对搜索关键词做URL编码,避免特殊字符导致请求错误
encoded_content = requests.utils.quote(content)
url = f"https://www.youtube.com/results?search_query={encoded_content}"

# 添加请求头获取真实页面内容
response = requests.get(url, headers=headers)
soup = BeautifulSoup(response.content, 'html.parser')

set_likes = set()
set_links = set()

# 精准定位搜索结果中的视频链接
video_links = soup.find_all("a", href=True, id="video-title")
for video in video_links:
    video_href = video['href']
    if "/watch?" in video_href:
        video_url = f"https://www.youtube.com{video_href}"
        # 再次添加请求头获取视频页内容
        video_response = requests.get(video_url, headers=headers)
        video_soup = BeautifulSoup(video_response.content, 'html.parser')
        
        # 适配当前YouTube页面结构的点赞数提取逻辑
        try:
            # 通过aria-label属性定位点赞按钮,提取其中的数字
            like_btn = video_soup.find("button", {"aria-label": lambda x: x and "j'aime" in x.lower()})
            if like_btn:
                likes = like_btn.get("aria-label").split()[0]
                set_likes.add(likes)
                set_links.add(video_url)
            else:
                print(f"无法获取视频点赞数: {video_url}")
        except Exception as e:
            print(f"处理视频时出错 {video_url}: {str(e)}")

# 输出结果
print("links:", set_links)
print("likes:", set_likes)

关键修复说明

  • 添加请求头:用User-Agent模拟浏览器请求,避免被YouTube直接拦截。
  • URL编码关键词:处理用户输入的特殊字符,确保请求URL合法。
  • 精准筛选视频链接:通过id="video-title"定位真实视频链接,减少无效请求。
  • 更新点赞数提取逻辑:利用aria-label属性定位点赞按钮,提取其中的数字内容,适配当前YouTube页面结构。

内容的提问来源于stack exchange,提问作者Fire35

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 14:46:13