You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何获取Google搜索关键词的结果页面总数?技术实现求助

解决方案:提取Google搜索结果显示的条目总数

原代码问题分析

你的代码存在几个关键问题导致无法实现目标:

  • 搜索词错误使用了“टंकलेखन”,而非目标的indieea,且未保留原需求的精确搜索双引号
  • 硬编码定位aria-label="Page 9"的分页链接,无法适配动态变化的搜索结果页数
  • 未处理跳转最后一页后提取提示文本中数字的逻辑

修正后的Playwright代码

import asyncio
import re
from playwright.async_api import async_playwright

async def main():
    # 带双引号的精确搜索词,对应原链接的%22indieea%22
    search_term = '"indieea"'
    url = f"https://www.google.com/search?q={search_term}"

    async with async_playwright() as pw:
        # 启动浏览器,可添加headless=False方便调试
        browser = await pw.chromium.launch(headless=True)
        page = await browser.new_page()
        # 加载页面,等待网络请求完成确保内容加载完整
        await page.goto(url, wait_until="networkidle")

        # 定位最后一页的链接(Google分页最后一页有aria-label="Last page")
        last_page_link = page.locator('a[aria-label="Last page"]')
        if await last_page_link.count() > 0:
            # 点击跳转到最后一页
            await last_page_link.click()
            # 等待页面加载完成
            await page.wait_for_load_state("networkidle")

            # 定位提示文本元素,匹配目标提示内容
            prompt_text = await page.locator('div:has-text("we have omitted some entries very similar to")').text_content()
            if prompt_text:
                # 提取提示中的数字
                match = re.search(r'the (\d+) already displayed', prompt_text)
                if match:
                    print(f"结果总数:{match.group(1)}")
                else:
                    print("未找到匹配的数字")
            else:
                print("未找到提示文本")
        else:
            # 如果没有分页,直接提取结果统计数字
            stats_text = await page.locator('#result-stats').text_content()
            if stats_text:
                match = re.search(r'About (\d+) results', stats_text)
                if match:
                    print(f"结果总数:{match.group(1)}")
                else:
                    print("未找到结果统计数字")

        await browser.close()

if __name__ == "__main__":
    asyncio.run(main())

关键说明

  • 精确搜索处理:保留搜索词的双引号,确保和原需求的搜索链接行为一致
  • 动态分页定位:通过aria-label="Last page"定位最后一页,避免硬编码页数
  • 文本提取逻辑:使用正则表达式从提示文本中提取数字,适配Google的页面结构
  • 加载策略:使用networkidle等待页面完全加载,避免因内容未加载导致元素定位失败

注意事项

  • Google可能会根据地区、Cookie等因素调整页面结构,若定位失败需检查元素选择器
  • 频繁请求可能触发反爬机制,可适当添加等待时间或配置浏览器上下文模拟真实用户行为

内容的提问来源于stack exchange,提问作者shantanuo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 03:05:02