You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium爬取谷歌搜索赞助链接失败问题求助

谷歌搜索赞助链接爬取异常问题

使用Selenium批量爬取约1000个关键词的谷歌搜索赞助条目链接时,前几个关键词能正常获取赞助URL,但后续明明存在赞助广告的关键词却无法爬取到结果。相关代码如下:

def get_ads_url(keywords):
    chrome_options = webdriver.ChromeOptions()
    chrome_options.add_argument('--headless')
    chrome_options.add_argument('--no-sandbox')
    chrome_options.add_argument('--disable-dev-shm-usage')

    driver = webdriver.Chrome(chromedriver_path, options= chrome_options)

    google_url = 'https://www.google.com/search?q={}'.format(keywords) 

    driver.get(google_url)
    pause = float( random.randint(1, 2) + random.randint(1, 5)/ 10)
    time.sleep(pause)

    soup = BeautifulSoup(driver.page_source,'lxml')
    sponsored_items = soup.select("div[data-text-ad='1'] span[role='text']")

    sponsored_list = []
    for item in sponsored_items:
        text = item.get_text(strip = True).split()
        if len(text) > 0:
            sponsored_list.append(text[0])

    final_result = list(set(sponsored_list))

    print(final_result)
    driver.quit()

    return final_result

示例

  • 关键词: nike
  • 谷歌搜索URL: https://www.google.com/search?q=nike
  • 预期结果:
    ['https://www.nike.com.tw', 
     'https://www.adidas.com.tw',
     'https://www.farfetch.com']
    

耐克搜索结果的赞助条目截图:
Nike的赞助条目

解决思路

1. 应对谷歌反爬拦截

  • 伪装真实浏览器:无头浏览器特征明显,添加参数模拟普通用户浏览器:
    chrome_options.add_argument('--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36')
    chrome_options.add_argument('--disable-blink-features=AutomationControlled')
    chrome_options.add_experimental_option("excludeSwitches", ["enable-automation"])
    chrome_options.add_experimental_option('useAutomationExtension', False)
    
  • 降低请求频率:当前1-2.5秒等待时间过短,延长至3-8秒随机范围:
    pause = random.uniform(3, 8)
    time.sleep(pause)
    
  • 复用浏览器实例:不要每次请求都重启浏览器,减少爬虫特征暴露。

2. 修复页面元素定位逻辑

谷歌广告DOM结构可能动态调整,原选择器失效,改用更稳定的定位方式:

  • 定位广告容器后提取真实链接(处理谷歌跳转URL):
    from urllib.parse import urlparse, parse_qs
    
    soup = BeautifulSoup(driver.page_source,'lxml')
    ad_containers = soup.select("div[data-ad='1']")
    sponsored_list = []
    for container in ad_containers:
        links = container.select("a")
        for link in links:
            href = link.get("href")
            if href:
                if "url?q=" in href:
                    parsed_url = urlparse(href)
                    real_url = parse_qs(parsed_url.query)['url'][0]
                    sponsored_list.append(real_url)
                elif href.startswith("http"):
                    sponsored_list.append(href)
    final_result = list(set(sponsored_list))
    
  • 使用显式等待:替代固定sleep,等待广告元素加载完成后再解析:
    from selenium.webdriver.support.ui import WebDriverWait
    from selenium.webdriver.support import expected_conditions as EC
    from selenium.webdriver.common.by import By
    
    driver.get(google_url)
    try:
        WebDriverWait(driver, 10).until(
            EC.presence_of_element_located((By.CSS_SELECTOR, "div[data-ad='1']"))
        )
    except:
        pass  # 无广告则跳过
    

3. 解决IP与会话限制

  • 使用代理IP:切换代理分散请求来源,避免IP被封禁:
    chrome_options.add_argument('--proxy-server=http://your-proxy-ip:port')
    
  • 复用Cookie会话:保存浏览器Cookie,模拟真实用户连续访问行为,减少新用户检测概率。

内容的提问来源于stack exchange,提问作者Jammy Wang

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 00:54:52