如何在Python Playwright中获取下载文件的字节流?
获取Playwright下载文件的字节流(无需中间保存)
你可以通过以下两种方式直接获取下载文件的字节流,跳过本地保存步骤:
方法1:使用create_read_stream()读取流
Playwright的Download对象提供了create_read_stream()方法,能直接获取文件的可读流,再读取为字节:
with page.expect_download() as download_info: page.get_by_role("button", name="Download PDF").click() download = download_info.value # 获取文件流并读取为字节 with download.create_read_stream() as stream: file_bytes = stream.read() # 现在file_bytes就是文件的字节流,可以直接传给云存储API
方法2:读取临时下载文件
Playwright会把下载的文件存放在临时路径中,你可以直接读取该路径的文件内容:
with page.expect_download() as download_info: page.get_by_role("button", name="Download PDF").click() download = download_info.value # 获取临时文件路径 temp_path = download.path() # 读取文件字节 with open(temp_path, "rb") as f: file_bytes = f.read() # 同样可以直接使用file_bytes调用云API
注意:临时文件会在download对象被垃圾回收后自动清理,所以要确保在使用完字节流前,download对象始终处于有效状态。
内容的提问来源于stack exchange,提问作者Jeroen Vermunt
相关产品推荐
相关产品推荐

