Selenium爬取YouTube搜索结果时出现‘invalid argument: 'url' must be a string’错误的解决求助
Hey there, let's get to the bottom of this error you're seeing. The problem boils down to one key issue: some entries in your links list aren't valid string URLs—they're None values. Here's why and how to fix it:
Why the Error Happens
Your current XPath //*[@id="video-title"] matches more than just video links. YouTube's search results include things like channel cards, playlists, and sometimes non-video elements that share the same video-title ID but don't have a valid href attribute. When you call i.get_attribute('href') on these elements, you get None instead of a string URL. Then when you try driver.get(x) with x = None, Selenium throws that "url must be a string" error.
Step-by-Step Fixes
1. Filter Out Invalid/None Links
First, we need to ensure only valid URLs make it into your links list. Add a check when collecting links to skip any None values and non-video URLs:
# Replace your links collection code with this links = [] for i in user_data: href = i.get_attribute('href') # Only add if href exists and points to a video watch page if href and "/watch?v=" in href: links.append(href)
Or use a concise list comprehension:
links = [i.get_attribute('href') for i in user_data if i.get_attribute('href') and "/watch?v=" in i.get_attribute('href')]
2. Use a Precise XPath to Target Only Videos
To avoid pulling in non-video elements altogether, refine your XPath to only select anchor tags with the video-title ID that link to YouTube watch pages:
# Replace your user_data line with this user_data = driver.find_elements_by_xpath('//a[@id="video-title" and contains(@href, "/watch?v=")]')
This ensures you're only grabbing actual video links right from the start.
3. Optional: Improve Video ID Extraction
Your current v_id = x.strip('https://www.youtube.com/watch?v=') might fail if the URL has extra parameters (like &t=10s). A more reliable approach uses Python's urllib.parse module:
from urllib.parse import urlparse, parse_qs # Inside your loop: parsed_url = urlparse(x) v_id = parse_qs(parsed_url.query)['v'][0]
Full Modified Code
Here's the complete fixed code incorporating all these changes:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from urllib.parse import urlparse, parse_qs chrome_path = r'C:\Windows\chromedriver.exe' driver = webdriver.Chrome(chrome_path) driver.get("https://www.youtube.com/results?search_query=python+course") # Target only video links directly user_data = driver.find_elements_by_xpath('//a[@id="video-title" and contains(@href, "/watch?v=")]') links = [] for i in user_data: href = i.get_attribute('href') if href: links.append(href) print(links) wait = WebDriverWait(driver, 10) for x in links: driver.get(x) # Extract video ID reliably parsed_url = urlparse(x) v_id = parse_qs(parsed_url.query)['v'][0] # Wait for video title to load v_title = wait.until(EC.presence_of_element_located( (By.CSS_SELECTOR,"h1.title yt-formatted-string"))).text print(f"Video ID: {v_id}, Title: {v_title}")
This should eliminate the "url must be a string" error and make your scraper more robust against YouTube's varying result elements.
内容的提问来源于stack exchange,提问作者Mark J.

