You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python+Selenium爬取订阅数<500的YouTube创作者并导出Excel?

实现思路:筛选YouTube订阅数少于500的创作者并导出至Excel

核心步骤拆解

  • 切换到创作者搜索结果页:当前代码搜索后默认显示视频内容,需要点击顶部的「Channels」标签,切换到创作者列表页面。
  • 提取创作者信息与订阅数:遍历每个创作者卡片,获取频道名称、订阅数等数据。注意订阅数的文本格式(比如「450 subscribers」或「1.2K subscribers」),需要统一转换为数字以便筛选。
  • 筛选符合条件的创作者:判断转换后的订阅数是否小于500,保留符合条件的条目。
  • 导出数据到Excel:用pandas或openpyxl库将筛选后的数据写入Excel文件,这两个库都能轻松实现表格导出。

关键代码补充与优化

首先安装依赖库:

pip install selenium pandas openpyxl

修改后的完整代码:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import pandas as pd
import time

def convert_subs_to_num(subs_text):
    # 处理订阅数字符串,转换为整数
    subs_text = subs_text.lower().replace('subscribers', '').strip()
    if 'k' in subs_text:
        return int(float(subs_text.replace('k', '')) * 1000)
    elif 'm' in subs_text:
        return int(float(subs_text.replace('m', '')) * 1000000)
    else:
        return int(subs_text.replace(',', '')) if subs_text.isdigit() else 0

driver = webdriver.Chrome() 
path = "https://youtube.com"
driver.get(path)
driver.maximize_window()

# 等待搜索框加载并输入关键词
wait = WebDriverWait(driver, 10)
search_box = wait.until(EC.presence_of_element_located((By.XPATH, '//input[@id="search"]')))
search_word = input("Enter the search keyword: ")
search_box.send_keys(search_word)

# 点击搜索按钮
search_btn = wait.until(EC.element_to_be_clickable((By.ID, 'search-icon-legacy')))
search_btn.click()

# 切换到Channels标签
time.sleep(2)  # 等待搜索结果加载
channels_tab = wait.until(EC.element_to_be_clickable((By.XPATH, '//div[@id="tabsContent"]//a[@title="Channels"]')))
channels_tab.click()

# 滚动加载更多创作者(可选,根据需要加载更多内容)
last_height = driver.execute_script("return document.documentElement.scrollHeight")
while True:
    driver.execute_script("window.scrollTo(0, document.documentElement.scrollHeight);")
    time.sleep(2)
    new_height = driver.execute_script("return document.documentElement.scrollHeight")
    if new_height == last_height:
        break
    last_height = new_height

# 提取创作者数据
creators = []
channel_cards = wait.until(EC.presence_of_all_elements_located((By.XPATH, '//ytd-channel-renderer')))
for card in channel_cards:
    try:
        channel_name = card.find_element(By.XPATH, './/yt-formatted-string[@id="channel-title"]').text
        subs_text = card.find_element(By.XPATH, './/yt-formatted-string[@id="subscriber-count"]').text
        subs_num = convert_subs_to_num(subs_text)
        channel_link = card.find_element(By.XPATH, './/a[@id="channel-link"]').get_attribute('href')
        
        if subs_num < 500:
            creators.append({
                "Channel Name": channel_name,
                "Subscribers": subs_num,
                "Channel Link": channel_link
            })
    except Exception as e:
        print(f"提取卡片信息出错: {e}")
        continue

# 导出到Excel
df = pd.DataFrame(creators)
df.to_excel("small_youtube_creators.xlsx", index=False)
print(f"已导出{len(creators)}个符合条件的创作者数据到Excel文件")

driver.quit()

重要说明

  • 避免绝对XPATH:原代码使用的绝对XPATH易受页面结构变化影响,改用相对XPATH或元素ID定位更稳定。
  • 显式等待替代time.sleep:用WebDriverWait等待元素加载,比固定时长的time.sleep更可靠。
  • 订阅数转换逻辑:处理了K、M量级的订阅数,确保筛选准确。
  • 滚动加载:如果需要获取更多创作者,滚动页面加载所有内容后再提取数据。

内容的提问来源于stack exchange,提问作者Aleksandra Milicevic

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 10:28:38