You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用browser-use实现步骤截图与全程录屏功能?

解决browser-use自动截图与录屏无输出的问题

问题描述

我正在构建一个使用browser-use执行任务的Python项目,需要实现每步自动截图、全程录屏功能,但按官方文档配置后,无法生成对应的截图与录屏文件。当前配置代码如下:

from browser_use import Agent, BrowserContext, BrowserProfile, BrowserSession
import os

recordings_dir = os.path.join(temp_dir, 'recordings')

# Configure the browser session
browser_session = BrowserSession(
    headless=False,
    viewport={'width': 1920, 'height': 1080},
    disable_web_security=True,
    ignore_https_errors=True,
    disable_features=['VizDisplayCompositor'],
)

# Define the browser profile with recording settings
browser_profile = BrowserProfile(
    record_video_dir=recording_dir,
    record_video_size={'width': 1920, 'height': 1080},
    save_recording_path=recording_dir,
    deterministic_rendering=True,
    disable_images=False,
    disable_javascript=False,
    disable_animations=True,
    page_load_timeout=15000,
)

# Set up the browser context with video recording settings
browser_context = BrowserContext(
    enable_recording=True,
    recordings_dir=recording_dir,
    # record_video_dir=recording_dir,
    # record_video_size={'width': 1920, 'height': 1080},
    viewport={'width': 1920, 'height': 1080},
)

self.agent = Agent(
    task="Browser automation task with screenshot and video capture",
    llm=llm,
    browser_session=browser_session,
    browser_profile=browser_profile,
    browser_context=browser_context
)

发现screenshots和video_recording字段存在但内容为空,试过Playwright方案但不适用于browser-use,看到官方web-ui仓库已实现该功能,却找不到具体实现方式,求解决方法。


解决步骤

1. 清理冗余配置,统一参数位置

browser-use的配置优先级中,BrowserContext的录屏设置会覆盖BrowserProfile的相关配置,且部分参数存在重复定义。建议将所有录屏/截图相关配置统一放到BrowserContext中,移除BrowserProfile里的重复项:

# 简化后的BrowserProfile(移除录屏相关冗余配置)
browser_profile = BrowserProfile(
    deterministic_rendering=True,
    disable_images=False,
    disable_javascript=False,
    disable_animations=True,
    page_load_timeout=15000,
)

# 完整的BrowserContext配置(包含截图+录屏)
browser_context = BrowserContext(
    enable_recording=True,
    recordings_dir=recording_dir,
    record_video_size={'width': 1920, 'height': 1080},
    viewport={'width': 1920, 'height': 1080},
    auto_screenshot=True,
    screenshot_dir=os.path.join(temp_dir, 'screenshots')
)

2. 确保目录存在且有写入权限

在初始化目录后,显式创建目录并验证权限,避免因目录不存在导致写入失败:

import os

recordings_dir = os.path.join(temp_dir, 'recordings')
screenshots_dir = os.path.join(temp_dir, 'screenshots')

# 确保目录存在,不存在则创建
os.makedirs(recordings_dir, exist_ok=True)
os.makedirs(screenshots_dir, exist_ok=True)

# 验证目录可写入
assert os.access(recordings_dir, os.W_OK), f"无写入权限:{recordings_dir}"
assert os.access(screenshots_dir, os.W_OK), f"无写入权限:{screenshots_dir}"

3. 绑定事件钩子手动触发截图

如果自动截图未生效,可通过Agent的事件钩子在步骤完成时手动触发截图:

def on_step_completed(step_result):
    # 获取当前页面实例
    page = self.agent.browser_session.get_current_page()
    if page:
        # 生成截图路径
        screenshot_path = os.path.join(screenshots_dir, f"step_{step_result.step_id}.png")
        # 执行截图
        page.screenshot(path=screenshot_path)
        # 将截图路径关联到结果字段
        step_result.screenshots.append(screenshot_path)

# 绑定步骤完成事件
self.agent.on("step_completed", on_step_completed)

4. 参考官方Web UI的核心实现逻辑

官方Web UI的录屏/截图功能,核心是通过扩展BrowserSession的生命周期钩子实现:

  • 录屏:在session_started事件中启动视频录制,在session_ended事件中保存录屏文件
  • 截图:在step_executed事件前后调用页面截图方法,自动关联到任务结果

你可以通过监听browser-use的内部事件,在对应时机调用浏览器原生API完成录制操作。

5. 调试验证录制状态

在代码中添加调试代码,检查录制配置和文件生成情况:

# 初始化后检查配置状态
print(f"录屏启用状态:{browser_context.enable_recording}")
print(f"录屏目录:{browser_context.recordings_dir}")
print(f"截图目录:{browser_context.screenshot_dir}")

# 任务执行后检查文件
self.agent.run()
print(f"生成的录屏文件:{os.listdir(recordings_dir)}")
print(f"生成的截图文件:{os.listdir(screenshots_dir)}")

内容的提问来源于stack exchange,提问作者Terry Windwalker

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.12 22:52:36