如何获取Google搜索关键词的结果页面总数?技术实现求助
解决方案:提取Google搜索结果显示的条目总数
原代码问题分析
你的代码存在几个关键问题导致无法实现目标:
- 搜索词错误使用了“टंकलेखन”,而非目标的
indieea,且未保留原需求的精确搜索双引号 - 硬编码定位
aria-label="Page 9"的分页链接,无法适配动态变化的搜索结果页数 - 未处理跳转最后一页后提取提示文本中数字的逻辑
修正后的Playwright代码
import asyncio import re from playwright.async_api import async_playwright async def main(): # 带双引号的精确搜索词,对应原链接的%22indieea%22 search_term = '"indieea"' url = f"https://www.google.com/search?q={search_term}" async with async_playwright() as pw: # 启动浏览器,可添加headless=False方便调试 browser = await pw.chromium.launch(headless=True) page = await browser.new_page() # 加载页面,等待网络请求完成确保内容加载完整 await page.goto(url, wait_until="networkidle") # 定位最后一页的链接(Google分页最后一页有aria-label="Last page") last_page_link = page.locator('a[aria-label="Last page"]') if await last_page_link.count() > 0: # 点击跳转到最后一页 await last_page_link.click() # 等待页面加载完成 await page.wait_for_load_state("networkidle") # 定位提示文本元素,匹配目标提示内容 prompt_text = await page.locator('div:has-text("we have omitted some entries very similar to")').text_content() if prompt_text: # 提取提示中的数字 match = re.search(r'the (\d+) already displayed', prompt_text) if match: print(f"结果总数:{match.group(1)}") else: print("未找到匹配的数字") else: print("未找到提示文本") else: # 如果没有分页,直接提取结果统计数字 stats_text = await page.locator('#result-stats').text_content() if stats_text: match = re.search(r'About (\d+) results', stats_text) if match: print(f"结果总数:{match.group(1)}") else: print("未找到结果统计数字") await browser.close() if __name__ == "__main__": asyncio.run(main())
关键说明
- 精确搜索处理:保留搜索词的双引号,确保和原需求的搜索链接行为一致
- 动态分页定位:通过
aria-label="Last page"定位最后一页,避免硬编码页数 - 文本提取逻辑:使用正则表达式从提示文本中提取数字,适配Google的页面结构
- 加载策略:使用
networkidle等待页面完全加载,避免因内容未加载导致元素定位失败
注意事项
- Google可能会根据地区、Cookie等因素调整页面结构,若定位失败需检查元素选择器
- 频繁请求可能触发反爬机制,可适当添加等待时间或配置浏览器上下文模拟真实用户行为
内容的提问来源于stack exchange,提问作者shantanuo
相关产品推荐
相关产品推荐

