使用Selenium爬取谷歌搜索结果报错:AttributeError: 'NoneType' object has no attribute 'text'
Fixing AttributeError When Scraping Google Search with Headless Selenium
Hey there! Let's break down why you're hitting that AttributeError: 'NoneType' object has no attribute 'text' error and get your Google search scraper working smoothly. That error means you're trying to call .text on a None object—basically, Selenium couldn't find the element you were targeting, usually because the page wasn't loaded yet or your selector is outdated.
Here's a Fixed, Working Version of Your Code
from selenium import webdriver from selenium.webdriver.common.keys import Keys from selenium.webdriver.chrome.options import Options from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By # Configure headless Chrome options chrome_options = Options() chrome_options.add_argument("--window-size=1024x768") chrome_options.add_argument("--headless=new") # Newer, more stable headless mode for Chrome chrome_options.add_argument("--disable-gpu") # Fixes rendering issues in headless mode chrome_options.add_argument("--no-sandbox") # Useful for Linux environments or restricted setups # Initialize the driver driver = webdriver.Chrome(options=chrome_options) try: # Navigate to Google Search driver.get("https://www.google.com") # Wait for the search box to load, then input your query search_box = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.NAME, "q")) ) search_box.send_keys("Your search query here") search_box.send_keys(Keys.RETURN) # Wait for search results to fully load result_containers = WebDriverWait(driver, 10).until( EC.presence_of_all_elements_located((By.CSS_SELECTOR, "div.g")) ) # Print out the results print("Google Search Results:") for count, container in enumerate(result_containers, 1): try: # Extract title and link from each result title = container.find_element(By.CSS_SELECTOR, "h3").text link = container.find_element(By.CSS_SELECTOR, "a").get_attribute("href") print(f"{count}. {title}\n{link}\n") except Exception: # Skip any problematic elements (like ads or non-standard results) continue finally: # Make sure the browser closes even if something goes wrong driver.quit()
Key Fixes & Explanations
- Wait for elements to load properly: Instead of grabbing elements immediately, we use
WebDriverWaitwith a 10-second timeout. This ensures Selenium only tries to access elements once they're actually present on the page—no moreNoneobjects from premature access. - Update headless mode: Chrome's old
--headlessflag has been replaced with--headless=newfor better stability and compatibility with modern Chrome versions. - Use reliable selectors: Google's search results are wrapped in
div.gcontainers, with titles in<h3>tags and links in<a>tags. These selectors are more stable than older ones that might have changed. - Add error handling: Wrapping the title/link extraction in a
try-exceptblock lets us skip any odd elements (like ads or featured snippets) that don't match our expected structure, preventing the whole script from crashing. - Clean up resources: The
finallyblock guarantees the browser closes even if the script hits an error, so you don't have leftover Chrome processes running in the background.
Quick Additional Tips
- Make sure your ChromeDriver version matches your installed Chrome version—mismatches cause all sorts of weird issues.
- If you're scraping from a region-specific Google domain (like
google.co.uk), update the URL indriver.get()accordingly. - For Linux servers, the
--no-sandboxflag is often necessary to avoid permission errors when running headless Chrome.
内容的提问来源于stack exchange,提问作者Big-karl Sisamu Sibalatani
相关产品推荐
相关产品推荐

