如何使用Selenium和Python对YouTube进行网络数据爬取
Hey there! Let's figure out how to scrape math formulas and symbols from YouTube using your existing Selenium code. I’ll walk you through refining it step by step.
First, let's fix a small syntax issue in your original code (Python uses # for comments, not //) and build out the parts needed to extract math-related formulas and symbols.
Step 1: Update Imports & Base Setup
We'll add the By module for modern element locating, and switch to HTTPS for a secure connection to YouTube:
from selenium import webdriver from selenium.webdriver.chrome.options import Options from selenium.webdriver.common.by import By import time import json options = Options() options.headless = False driver = webdriver.Chrome(options=options) # Fixed comment syntax driver.implicitly_wait(5) baseurl = 'https://youtube.com' keyword = input("Enter a math-related keyword (e.g., 'calculus formulas'): ") driver.get(f'{baseurl}/search?q={keyword}')
Step 2: Extract Math-Focused Content
YouTube loads results dynamically, so we’ll wait a moment for everything to render, then pull video titles, descriptions, and URLs that contain math symbols or formulas. We’ll filter entries to only keep those with recognizable math characters:
# Wait for dynamic content to fully load (implicit wait might not catch all) time.sleep(3) # Locate all video titles and descriptions video_titles = driver.find_elements(By.ID, 'video-title') video_descriptions = driver.find_elements(By.CSS_SELECTOR, '#description-text') # Store our extracted math content math_entries = [] # Loop through each result and filter for math-related content for title, desc in zip(video_titles, video_descriptions): title_text = title.text.strip() desc_text = desc.text.strip() # Check for common math symbols (expand this set as needed!) math_symbols = {'=', '+', '−', '×', '÷', '∑', 'π', '√', '∫', '∞', '∈', '∼'} if any(symbol in title_text + desc_text for symbol in math_symbols): math_entries.append({ 'video_title': title_text, 'video_url': title.get_attribute('href'), 'description': desc_text }) # Save the results to a JSON file for easy access with open('youtube_math_content.json', 'w', encoding='utf-8') as f: json.dump(math_entries, f, indent=2, ensure_ascii=False) print(f"Successfully extracted {len(math_entries)} math-related entries!") # Always close the driver when done driver.quit()
Step 3: Bonus Tips for Better Scraping
- Load More Results: YouTube only shows a handful of results initially. Add this scrolling loop to load more content:
# Scroll to load additional results last_scroll_height = driver.execute_script("return document.documentElement.scrollHeight") while True: driver.execute_script("window.scrollTo(0, document.documentElement.scrollHeight);") time.sleep(2) new_scroll_height = driver.execute_script("return document.documentElement.scrollHeight") if new_scroll_height == last_scroll_height: break last_scroll_height = new_scroll_height - Advanced Formula Detection: For more complex formulas, use regex to match patterns like
r'\b\d+[\+\-\×÷]\d+\b'(for basic arithmetic) orr'∫\w+\s*d\w+'(for integrals). - Respect YouTube's Rules: Make sure to avoid sending too many requests too quickly—this can get your IP blocked. Stick to reasonable delays between actions.
内容的提问来源于stack exchange,提问作者Kavita Negi

