基于Selenium的YouTube创作者爬取:求订阅量<500的筛选方案
解决方案:用Selenium筛选订阅量少于500的YouTube创作者
核心问题修正与步骤说明
你的原始代码收集了页面所有<a>标签的链接,但其中包含大量非视频链接(比如导航栏、搜索结果筛选栏链接),首先需要过滤出有效视频链接;其次要针对每个视频页面提取创作者的订阅量,判断是否符合小于500的条件。
修改后的完整代码
from selenium import webdriver import time from selenium.webdriver.common.by import By from selenium.webdriver.chrome.options import Options import pandas as pd from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC options = webdriver.ChromeOptions() options.add_experimental_option("detach", True) options.add_argument("--disable-extensions") options.add_argument("--disable-notifications") options.add_argument("--disable-popup-blocking") options.add_argument("start-maximized") driver= webdriver.Chrome(options=options) driver.implicitly_wait(10) # 全局隐式等待,提升稳定性 # 访问YouTube主页并搜索关键词 path = "https://youtube.com" driver.get(path) # 定位搜索框并输入关键词 search_box = driver.find_element(By.XPATH, '//input[@id="search"]') # 改用更稳定的ID定位,避免绝对XPATH失效 search_word = input("Enter the search keyword: ") search_box.send_keys(search_word) # 点击搜索按钮 search_button = driver.find_element(By.ID, "search-icon-legacy") search_button.click() # 滚动加载搜索结果(最多加载10次滚动) SCROLL_PAUSE_TIME = 2 last_height = driver.execute_script("return document.documentElement.scrollHeight") scroll_count = 0 while scroll_count < 10: driver.execute_script("window.scrollTo(0, document.documentElement.scrollHeight);") time.sleep(SCROLL_PAUSE_TIME) new_height = driver.execute_script("return document.documentElement.scrollHeight") if new_height == last_height: break last_height = new_height scroll_count += 1 # 过滤出有效视频链接(包含/watch?v=且去重) video_links = [] all_links = driver.find_elements(By.TAG_NAME, "a") for link in all_links: href = link.get_attribute("href") if href and "/watch?v=" in href and href not in video_links: video_links.append(href) # 存储符合条件的创作者信息 qualified_creators = [] # 遍历每个视频链接,提取创作者订阅量 for link in video_links[:10]: # 先测试前10个链接,可根据需求调整数量 try: driver.get(link) # 等待创作者名称元素加载 creator_name = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.XPATH, '//ytd-channel-name//a')) ).text # 等待订阅量元素加载 subscriber_count_elem = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.ID, 'subscriber-count')) ) subscriber_text = subscriber_count_elem.text # 解析订阅量数字:处理"X subscribers"或"XK subscribers"格式 if "K" in subscriber_text: # 订阅量以千为单位,直接跳过(因为大于500) continue # 提取数字部分 subscriber_num = int(subscriber_text.split()[0].replace(',', '')) if subscriber_num < 500: qualified_creators.append({ "创作者名称": creator_name, "订阅量": subscriber_num, "视频链接": link }) print(f"找到符合条件的创作者:{creator_name}({subscriber_num}订阅)") except Exception as e: print(f"处理链接{link}时出错:{str(e)}") continue # 将结果保存为CSV if qualified_creators: df = pd.DataFrame(qualified_creators) df.to_csv("qualified_youtube_creators.csv", index=False, encoding='utf-8-sig') print("结果已保存到qualified_youtube_creators.csv") else: print("未找到订阅量少于500的创作者") driver.quit()
关键修改点说明
- 链接过滤:只保留包含
/watch?v=的链接并去重,避免无效链接干扰。 - 元素定位优化:改用ID或相对XPATH定位元素,避免绝对XPATH因页面结构变化失效。
- 显式等待:使用
WebDriverWait替代time.sleep,确保元素加载完成后再操作,提升稳定性。 - 订阅量解析:处理常见的订阅量文本格式,自动过滤掉以"K"为单位的账号(订阅量≥1000),只保留数字小于500的创作者。
- 异常处理:捕获单个视频处理时的错误,避免程序中断。
注意事项
- YouTube页面结构可能随时更新,如果元素定位失效,需要根据当前页面的HTML结构调整XPATH或ID。
- 频繁访问可能触发YouTube的反爬机制,建议适当增加等待时间,避免短时间内大量请求。
内容的提问来源于stack exchange,提问作者Aleksandra Milicevic
相关产品推荐
相关产品推荐

