You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Playwright-Python异步爬虫中page.close()无法正常关闭页面的问题

问题分析与解决方案

你的核心问题有两个:

  1. 当前代码中仅写了page.close,未实际调用方法,且Playwright异步API的page.close()是异步方法,必须用await关键字执行
  2. 未确保页面在爬取完成(或出错)时被关闭,导致大量页面实例堆积占用内存

正确的页面关闭方式

方案1:用try/finally确保页面必关

在爬取逻辑外包裹try/finally块,无论爬取过程是否出错,都能保证页面被关闭:

async def scrape(context, url):
    page = await context.new_page()
    try:
        await page.goto(url) 
        await page.wait_for_load_state(state="networkidle")
        await page.wait_for_timeout(1000)
        # 获取数据逻辑不变
        html = await page.content()
        soup = BeautifulSoup(html, "lxml")
        tables = soup.find_all('table')
        dfs = pd.read_html(str(tables))
        df=dfs[0]
        print(f"Dataframe in page {url} scraped")
        return df
    finally:
        await page.close()  # 异步关闭页面,必须加await和括号

方案2:用async with自动管理页面生命周期

Playwright支持用async with创建页面,代码块结束时会自动关闭页面,更简洁安全:

async def scrape(context, url):
    async with context.new_page() as page:  # 自动在代码块结束后关闭页面
        await page.goto(url) 
        await page.wait_for_load_state(state="networkidle")
        await page.wait_for_timeout(1000)
        html = await page.content()
        soup = BeautifulSoup(html, "lxml")
        tables = soup.find_all('table')
        dfs = pd.read_html(str(tables))
        df=dfs[0]
        print(f"Dataframe in page {url} scraped")
        return df

补充:优化浏览器上下文与资源管理

为进一步避免内存泄漏,建议在main函数中用async with管理浏览器和上下文,自动释放资源:

async def main(urls):
    async with async_playwright() as p:
        async with p.chromium.launch(headless=False) as browser:
            async with browser.new_context() as context:
                master_results = pd.DataFrame()
                async with aiometer.amap(
                    functools.partial(scrape, context),
                    urls,
                    max_at_once=5, # 限制并发数
                    max_per_second=3,  # 限制请求频率
                ) as results:
                    async for data in results:
                        print(data)
                        master_results = pd.concat([master_results,data], ignore_index=True)
                print(master_results)

关于你遇到的TypeError

你之前调用await page.close()报错,大概率是写法错误(比如写成了await page.close,缺少括号),在异步函数中await page.close()是完全合法的异步调用,不会触发该错误。

内容的提问来源于stack exchange,提问作者NanoNerd

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 08:10:38