You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python+Playwright搜索自动化脚本故障排查与需求适配

问题分析与修复方案

核心问题定位

  1. 'NoneType' object is not subscriptable错误:
    出现在HumanBehavior.human_scroll方法中,page.viewport_size可能返回None(比如页面未完全加载、viewport未初始化时),直接访问['height']就会触发该错误。

  2. 脚本卡住并循环打开空白页:

    • 点击搜索结果链接时默认在当前标签页跳转,后续page.go_back()虽能回到搜索页,但频繁跳转易导致页面状态异常,触发超时或元素找不到,进而进入错误分支重新打开Google主页,形成循环。
    • 未处理新标签页场景:部分搜索结果会强制打开新标签页,脚本仍在原页面操作,导致后续步骤全部失效。
  3. 未实现7-10分钟的会话时长控制:
    原代码未统计整个搜索会话的时长,无法保证操作持续7-10分钟后再进入新循环。

修复后的完整代码

import random
import asyncio
from playwright.async_api import async_playwright

class HumanBehavior:
    @staticmethod
    async def random_delay(a=1, b=5):
        base = random.uniform(a, b)
        await asyncio.sleep(base * (0.8 + random.random() * 0.4))

    @staticmethod
    async def human_type(page, selector, text):
        for char in text:
            await page.type(selector, char, delay=random.randint(50, 200))
            if random.random() < 0.07:
                await page.keyboard.press('Backspace')
                await HumanBehavior.random_delay(0.1, 0.3)
                await page.type(selector, char)
            if random.random() < 0.2 and char == ' ':
                await HumanBehavior.random_delay(0.2, 0.5)

    @staticmethod
    async def human_scroll(page):
        # 改用evaluate获取窗口高度,避免viewport_size为None的问题
        viewport_height = await page.evaluate('window.innerHeight')
        for _ in range(random.randint(3, 7)):
            scroll_distance = random.randint(
                int(viewport_height * 0.5), 
                int(viewport_height * 1.5)
            )
            if random.random() < 0.3:
                scroll_distance *= -1
            await page.mouse.wheel(0, scroll_distance)
            await HumanBehavior.random_delay(0.7, 2.3)

    @staticmethod
    async def handle_popups(page):
        popup_selectors = [
            ('button:has-text("Accept")', 0.7),
            ('div[aria-label="Close"]', 0.5),
            ('button.close', 0.3),
            ('div.cookie-banner button', 0.4)  # 优化选择器,避免误点整个banner
        ]
        for selector, prob in popup_selectors:
            try:
                if random.random() < prob and await page.is_visible(selector, timeout=3000):
                    await page.click(selector)
                    await HumanBehavior.random_delay(0.5, 1.2)
            except Exception:
                continue  # 忽略弹窗处理失败的情况

async def perform_search_session(page):
    try:
        theme = "mental health"
        modifiers = ["how to", "best ways to", "guide for", "tips for"]
        query = f"{random.choice(modifiers)} {theme}"
        await page.goto("https://www.google.com", timeout=60000)
        await HumanBehavior.random_delay(2, 4)
        await HumanBehavior.handle_popups(page)
        
        # 等待搜索框并输入
        await page.wait_for_selector('textarea[name="q"]', timeout=10000)
        await HumanBehavior.human_type(page, 'textarea[name="q"]', query)
        await HumanBehavior.random_delay(0.5, 1.5)
        await page.keyboard.press('Enter')
        
        # 等待搜索结果加载完成
        await page.wait_for_selector('div.g', timeout=15000)
        await HumanBehavior.random_delay(2, 4)
        
        # 获取搜索结果链接(仅保留前10个有效结果)
        results = await page.query_selector_all('div.g a')
        results = [result for result in results if await result.get_attribute('href') and not await result.get_attribute('href').startswith('#')]
        if not results:
            print("No valid search results found")
            return False
        
        # 控制会话总时长在7-10分钟(420-600秒)
        session_start_time = asyncio.get_event_loop().time()
        target_duration = random.randint(420, 600)
        pages_opened = 0
        max_pages = random.randint(3, 7)
        
        while pages_opened < max_pages and (asyncio.get_event_loop().time() - session_start_time) < target_duration:
            # 随机选择前5个结果中的一个
            link = random.choice(results[:min(5, len(results))])
            link_href = await link.get_attribute('href')
            
            # 打开新标签页访问链接,避免干扰原搜索页
            async with page.context.new_page() as new_page:
                await new_page.goto(link_href, timeout=30000)
                await new_page.wait_for_load_state('networkidle', timeout=20000)
                await HumanBehavior.random_delay(3, 6)
                
                # 页面交互
                await HumanBehavior.human_scroll(new_page)
                await HumanBehavior.handle_popups(new_page)
                
                # 随机点击内部链接
                internal_links = await new_page.query_selector_all('a')
                internal_links = [link for link in internal_links if await link.get_attribute('href') and not await link.get_attribute('href').startswith('#')]
                if internal_links:
                    clicks = random.randint(1, 3)
                    for _ in range(clicks):
                        internal_link = random.choice(internal_links[:10])
                        internal_href = await internal_link.get_attribute('href')
                        await new_page.goto(internal_href, timeout=20000)
                        await new_page.wait_for_load_state('networkidle', timeout=20000)
                        await HumanBehavior.random_delay(2, 5)
                        await HumanBehavior.human_scroll(new_page)
                        await new_page.go_back(timeout=15000)
                        await HumanBehavior.random_delay(1, 3)
            
            pages_opened += 1
            await HumanBehavior.random_delay(2, 4)
            
            # 回到搜索页后重新获取结果(防止页面状态变化)
            await page.wait_for_selector('div.g', timeout=15000)
            results = await page.query_selector_all('div.g a')
            results = [result for result in results if await result.get_attribute('href') and not await result.get_attribute('href').startswith('#')]
            if not results:
                break
        
        # 如果时长未达标,继续在搜索页滚动或等待
        remaining_time = target_duration - (asyncio.get_event_loop().time() - session_start_time)
        if remaining_time > 0:
            await HumanBehavior.human_scroll(page)
            await asyncio.sleep(remaining_time)
        
        return True
    except Exception as e:
        print(f"Search error: {str(e)}")
        return False

# 示例运行代码
async def main():
    async with async_playwright() as p:
        browser = await p.chromium.launch(headless=False)
        context = await browser.new_context(viewport={'width': 1920, 'height': 1080})
        page = await context.new_page()
        while True:
            success = await perform_search_session(page)
            if success:
                print("Session completed, starting new search in 5-10 minutes...")
                await asyncio.sleep(random.randint(300, 600))
            else:
                print("Session failed, retrying in 2-5 minutes...")
                await asyncio.sleep(random.randint(120, 300))

if __name__ == "__main__":
    asyncio.run(main())

关键修复说明

  1. 解决NoneType下标错误:
    把page.viewport_size['height']替换为await page.evaluate('window.innerHeight'),直接从浏览器获取窗口高度,避免viewport未初始化的问题。

  2. 解决脚本卡住问题:

    • 使用page.context.new_page()打开新标签页访问搜索结果,原搜索页保持在后台,避免跳转导致的状态混乱。
    • 过滤无效链接(比如锚点链接),确保点击的是真实外部链接。
    • 增强弹窗处理的异常捕获,避免弹窗操作失败中断整个流程。
  3. 实现7-10分钟会话时长控制:

    • 记录会话开始时间,设置目标时长420-600秒。
    • 循环打开结果页时同时检查时长,若提前完成页面操作,通过滚动和等待补足剩余时间。
  4. 其他稳定性优化:

    • 增加链接有效性校验,过滤空链接和锚点链接。
    • 延长部分操作的超时时间,避免网络波动导致的失败。
    • 新增主循环逻辑,实现搜索会话完成后等待5-10分钟再进入下一轮循环。

内容的提问来源于stack exchange,提问作者Matthew

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 12:05:53