如何用browser-use实现步骤截图与全程录屏功能?
解决browser-use自动截图与录屏无输出的问题
问题描述
我正在构建一个使用browser-use执行任务的Python项目,需要实现每步自动截图、全程录屏功能,但按官方文档配置后,无法生成对应的截图与录屏文件。当前配置代码如下:
from browser_use import Agent, BrowserContext, BrowserProfile, BrowserSession import os recordings_dir = os.path.join(temp_dir, 'recordings') # Configure the browser session browser_session = BrowserSession( headless=False, viewport={'width': 1920, 'height': 1080}, disable_web_security=True, ignore_https_errors=True, disable_features=['VizDisplayCompositor'], ) # Define the browser profile with recording settings browser_profile = BrowserProfile( record_video_dir=recording_dir, record_video_size={'width': 1920, 'height': 1080}, save_recording_path=recording_dir, deterministic_rendering=True, disable_images=False, disable_javascript=False, disable_animations=True, page_load_timeout=15000, ) # Set up the browser context with video recording settings browser_context = BrowserContext( enable_recording=True, recordings_dir=recording_dir, # record_video_dir=recording_dir, # record_video_size={'width': 1920, 'height': 1080}, viewport={'width': 1920, 'height': 1080}, ) self.agent = Agent( task="Browser automation task with screenshot and video capture", llm=llm, browser_session=browser_session, browser_profile=browser_profile, browser_context=browser_context )
发现screenshots和video_recording字段存在但内容为空,试过Playwright方案但不适用于browser-use,看到官方web-ui仓库已实现该功能,却找不到具体实现方式,求解决方法。
解决步骤
1. 清理冗余配置,统一参数位置
browser-use的配置优先级中,BrowserContext的录屏设置会覆盖BrowserProfile的相关配置,且部分参数存在重复定义。建议将所有录屏/截图相关配置统一放到BrowserContext中,移除BrowserProfile里的重复项:
# 简化后的BrowserProfile(移除录屏相关冗余配置) browser_profile = BrowserProfile( deterministic_rendering=True, disable_images=False, disable_javascript=False, disable_animations=True, page_load_timeout=15000, ) # 完整的BrowserContext配置(包含截图+录屏) browser_context = BrowserContext( enable_recording=True, recordings_dir=recording_dir, record_video_size={'width': 1920, 'height': 1080}, viewport={'width': 1920, 'height': 1080}, auto_screenshot=True, screenshot_dir=os.path.join(temp_dir, 'screenshots') )
2. 确保目录存在且有写入权限
在初始化目录后,显式创建目录并验证权限,避免因目录不存在导致写入失败:
import os recordings_dir = os.path.join(temp_dir, 'recordings') screenshots_dir = os.path.join(temp_dir, 'screenshots') # 确保目录存在,不存在则创建 os.makedirs(recordings_dir, exist_ok=True) os.makedirs(screenshots_dir, exist_ok=True) # 验证目录可写入 assert os.access(recordings_dir, os.W_OK), f"无写入权限:{recordings_dir}" assert os.access(screenshots_dir, os.W_OK), f"无写入权限:{screenshots_dir}"
3. 绑定事件钩子手动触发截图
如果自动截图未生效,可通过Agent的事件钩子在步骤完成时手动触发截图:
def on_step_completed(step_result): # 获取当前页面实例 page = self.agent.browser_session.get_current_page() if page: # 生成截图路径 screenshot_path = os.path.join(screenshots_dir, f"step_{step_result.step_id}.png") # 执行截图 page.screenshot(path=screenshot_path) # 将截图路径关联到结果字段 step_result.screenshots.append(screenshot_path) # 绑定步骤完成事件 self.agent.on("step_completed", on_step_completed)
4. 参考官方Web UI的核心实现逻辑
官方Web UI的录屏/截图功能,核心是通过扩展BrowserSession的生命周期钩子实现:
- 录屏:在
session_started事件中启动视频录制,在session_ended事件中保存录屏文件 - 截图:在
step_executed事件前后调用页面截图方法,自动关联到任务结果
你可以通过监听browser-use的内部事件,在对应时机调用浏览器原生API完成录制操作。
5. 调试验证录制状态
在代码中添加调试代码,检查录制配置和文件生成情况:
# 初始化后检查配置状态 print(f"录屏启用状态:{browser_context.enable_recording}") print(f"录屏目录:{browser_context.recordings_dir}") print(f"截图目录:{browser_context.screenshot_dir}") # 任务执行后检查文件 self.agent.run() print(f"生成的录屏文件:{os.listdir(recordings_dir)}") print(f"生成的截图文件:{os.listdir(screenshots_dir)}")
内容的提问来源于stack exchange,提问作者Terry Windwalker
相关产品推荐
相关产品推荐

