You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup4提取谷歌搜索首图源链接遇TypeError错误求助

Fixing the "TypeError: 'NoneType' object is not subscriptable" in Your Google Image Scraper

Hey there! Let's break down why you're running into this error and how to fix it.

What's Causing the Error?

The error pops up because soup.find('img', class_='irc_mi') is returning None—meaning BeautifulSoup couldn't find any <img> tag with the class irc_mi in the page source. When you try to access ['src'] on None, Python throws that TypeError.

Here are the most common reasons this happens:

  • Google updated its page structure: The class name irc_mi is outdated—Google regularly changes HTML classes to prevent scraping.
  • Missing browser-like headers: Without proper request headers, Google might serve you a simplified or anti-bot page that doesn't include the image elements you're looking for.
  • Static content limits: The requests library only fetches the initial HTML, but some Google images load dynamically with JavaScript—so they won't show up in the static source.

How to Fix It

Here are practical solutions to get your scraper working again:

1. Update Your Request & Use Current HTML Selectors

First, mimic a real browser to avoid being blocked, and use the latest class name for Google Images (you'll need to verify this yourself):

import requests
from bs4 import BeautifulSoup

def get_first_google_image_link(theurl):
    # Mimic a Chrome browser request to avoid anti-bot measures
    headers = {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
        'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8'
    }

    try:
        r = requests.get(theurl, headers=headers)
        r.raise_for_status()  # Catch HTTP errors like 403 Forbidden
        soup = BeautifulSoup(r.text, "lxml")

        # Important: Update this class name! Right-click the first Google Image → Inspect to find the current one
        img_element = soup.find('img', class_='YQ4gaf')  # Example class name (may change soon!)

        if img_element:
            # Lazy-loaded images often store the real URL in data-src instead of src
            image_link = img_element.get('data-src') or img_element.get('src')
            return image_link
        else:
            return "Couldn't find the image element—double-check the class name in Google's HTML."
    except Exception as e:
        return f"Error occurred: {str(e)}"

2. Handle Dynamic Content with Selenium (If Needed)

If requests still can't find the image (because it's loaded via JavaScript), use Selenium to simulate a real browser that executes JS:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

def get_first_google_image_link_selenium(theurl):
    # Configure Chrome to run in headless mode (no visible window)
    chrome_options = Options()
    chrome_options.add_argument('--headless=new')
    chrome_options.add_argument('--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36')

    driver = webdriver.Chrome(options=chrome_options)
    try:
        driver.get(theurl)
        # Wait up to 10 seconds for the first image to load
        img_element = WebDriverWait(driver, 10).until(
            EC.presence_of_element_located((By.CSS_SELECTOR, 'img.YQ4gaf'))
        )
        # Grab the real image URL (check data-src first for lazy-loaded images)
        image_link = img_element.get_attribute('data-src') or img_element.get_attribute('src')
        return image_link
    except Exception as e:
        return f"Error occurred: {str(e)}"
    finally:
        driver.quit()  # Always close the browser when done

Quick Reminders

  • Google's classes change often: Always inspect the current HTML of Google Images to get the correct class name for image tags. Right-click the first image → Inspect to find the latest selector.
  • Respect Google's Terms of Service: Scraping Google may violate their terms, so make sure you're using this for personal, non-commercial use only.

内容的提问来源于stack exchange,提问作者Naveen Manoharan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:37:14