You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Selenium的YouTube创作者爬取:求订阅量<500的筛选方案

解决方案:用Selenium筛选订阅量少于500的YouTube创作者

核心问题修正与步骤说明

你的原始代码收集了页面所有<a>标签的链接,但其中包含大量非视频链接(比如导航栏、搜索结果筛选栏链接),首先需要过滤出有效视频链接;其次要针对每个视频页面提取创作者的订阅量,判断是否符合小于500的条件。

修改后的完整代码

from selenium import webdriver
import time
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.options import Options
import pandas as pd
from selenium.webdriver.support.ui import WebDriverWait 
from selenium.webdriver.support import expected_conditions as EC

options = webdriver.ChromeOptions()
options.add_experimental_option("detach", True)
options.add_argument("--disable-extensions")
options.add_argument("--disable-notifications")
options.add_argument("--disable-popup-blocking")
options.add_argument("start-maximized")

driver= webdriver.Chrome(options=options)
driver.implicitly_wait(10)  # 全局隐式等待,提升稳定性

# 访问YouTube主页并搜索关键词
path = "https://youtube.com"
driver.get(path)

# 定位搜索框并输入关键词
search_box = driver.find_element(By.XPATH, '//input[@id="search"]')  # 改用更稳定的ID定位,避免绝对XPATH失效
search_word = input("Enter the search keyword: ")
search_box.send_keys(search_word)

# 点击搜索按钮
search_button = driver.find_element(By.ID, "search-icon-legacy")
search_button.click()

# 滚动加载搜索结果(最多加载10次滚动)
SCROLL_PAUSE_TIME = 2
last_height = driver.execute_script("return document.documentElement.scrollHeight")
scroll_count = 0

while scroll_count < 10:
    driver.execute_script("window.scrollTo(0, document.documentElement.scrollHeight);")
    time.sleep(SCROLL_PAUSE_TIME)
    new_height = driver.execute_script("return document.documentElement.scrollHeight")
    if new_height == last_height:
        break
    last_height = new_height
    scroll_count += 1

# 过滤出有效视频链接(包含/watch?v=且去重)
video_links = []
all_links = driver.find_elements(By.TAG_NAME, "a")
for link in all_links:
    href = link.get_attribute("href")
    if href and "/watch?v=" in href and href not in video_links:
        video_links.append(href)

# 存储符合条件的创作者信息
qualified_creators = []

# 遍历每个视频链接,提取创作者订阅量
for link in video_links[:10]:  # 先测试前10个链接,可根据需求调整数量
    try:
        driver.get(link)
        # 等待创作者名称元素加载
        creator_name = WebDriverWait(driver, 10).until(
            EC.presence_of_element_located((By.XPATH, '//ytd-channel-name//a'))
        ).text
        # 等待订阅量元素加载
        subscriber_count_elem = WebDriverWait(driver, 10).until(
            EC.presence_of_element_located((By.ID, 'subscriber-count'))
        )
        subscriber_text = subscriber_count_elem.text
        
        # 解析订阅量数字:处理"X subscribers"或"XK subscribers"格式
        if "K" in subscriber_text:
            # 订阅量以千为单位,直接跳过(因为大于500)
            continue
        # 提取数字部分
        subscriber_num = int(subscriber_text.split()[0].replace(',', ''))
        
        if subscriber_num < 500:
            qualified_creators.append({
                "创作者名称": creator_name,
                "订阅量": subscriber_num,
                "视频链接": link
            })
            print(f"找到符合条件的创作者:{creator_name}({subscriber_num}订阅)")
    except Exception as e:
        print(f"处理链接{link}时出错:{str(e)}")
        continue

# 将结果保存为CSV
if qualified_creators:
    df = pd.DataFrame(qualified_creators)
    df.to_csv("qualified_youtube_creators.csv", index=False, encoding='utf-8-sig')
    print("结果已保存到qualified_youtube_creators.csv")
else:
    print("未找到订阅量少于500的创作者")

driver.quit()

关键修改点说明

  • 链接过滤:只保留包含/watch?v=的链接并去重,避免无效链接干扰。
  • 元素定位优化:改用ID或相对XPATH定位元素,避免绝对XPATH因页面结构变化失效。
  • 显式等待:使用WebDriverWait替代time.sleep,确保元素加载完成后再操作,提升稳定性。
  • 订阅量解析:处理常见的订阅量文本格式,自动过滤掉以"K"为单位的账号(订阅量≥1000),只保留数字小于500的创作者。
  • 异常处理:捕获单个视频处理时的错误,避免程序中断。

注意事项

  • YouTube页面结构可能随时更新,如果元素定位失效,需要根据当前页面的HTML结构调整XPATH或ID。
  • 频繁访问可能触发YouTube的反爬机制,建议适当增加等待时间,避免短时间内大量请求。

内容的提问来源于stack exchange,提问作者Aleksandra Milicevic

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 19:45:06