Python Playwright爬虫遇UnicodeDecodeError:如何识别响应类型并处理?
Playwright爬虫:响应类型判断及UnicodeDecodeError解决
问题场景
刚上手Python+Playwright开发网页爬虫,运行代码时触发UnicodeDecodeError,需要明确如何判断响应数据类型(文本/二进制等),并解决将响应转文本保存的报错问题。
原代码
from playwright.sync_api import sync_playwright import json def handle_response(response): with open("copy.txt", "w", encoding="utf-8") as file: file.write(response.text()) def main(): playwright=sync_playwright().start() browser=playwright.chromium.launch(headless=True) browser.new_context(no_viewport=True) page=browser.new_page() page.on('response',lambda response:handle_response(response)) page.goto("https://www.booking.com/hotel/it/hotelnordroma.en-gb.html?aid=304142&checkin=2025-05-15&checkout=2025-05-16#map_opened-map_trigger_header_pin") page.wait_for_timeout(1000) browser.close() playwright.stop() if __name__=='__main__': main()
报错信息
Exception has occurred: UnicodeDecodeError
'utf-8' codec can't decode byte 0x89 in position 0: invalid start byte
File "J:\SeSa\Playwright\sample.py", line 6, in handle_response
file.write(response.text())File "J:\SeSa\Playwright\sample.py", line 15, in page.on('response',lambda response:handle_response(response))<br /> ~~~~~~~~~~~~~~~~^^^^^^^^^^ UnicodeDecodeError: 'utf-8' codec can't decode byte 0x89 in position 0: invalid start byte
问题分析
报错里的0x89是PNG图片的起始字节,说明代码尝试将二进制响应(比如图片、字体、视频等非文本资源)强制解码为UTF-8文本,自然会触发解码失败。
解决方案
1. 如何判断响应数据类型
- 通过
response.content_type属性:直接获取响应的MIME类型,比如text/html(网页)、application/json(接口数据)属于文本类型;image/png(图片)、application/octet-stream(通用二进制)属于二进制类型。 - 通过响应头
Content-Type:和上面等价,可通过response.headers.get("Content-Type")获取。 - 方法区分:
response.text()会自动按响应编码解码为字符串,仅适用于文本资源;response.body()返回原始二进制字节,适用于所有类型的响应。
2. 修改代码解决报错
核心逻辑:只处理文本类型的响应,跳过或单独处理二进制资源。以下是优化后的代码:
from playwright.sync_api import sync_playwright def handle_response(response): content_type = response.content_type # 筛选需要处理的文本类型响应(可根据需求调整类型范围) text_types = ("text/", "application/json", "application/javascript") if any(content_type.startswith(t) for t in text_types): try: text_content = response.text() # 用追加模式写入,避免覆盖之前的内容 with open("copy.txt", "a", encoding="utf-8") as file: file.write(f"=== 响应URL: {response.url} ===\n") file.write(text_content + "\n\n") except Exception as e: print(f"处理文本响应失败[{response.url}]: {str(e)}") else: # 二进制资源可选择跳过,或保存为二进制文件 print(f"跳过二进制响应[{response.url}]: {content_type}") # 示例:保存图片类二进制资源 # if content_type.startswith("image/"): # filename = response.url.split("/")[-1] # with open(filename, "wb") as img_file: # img_file.write(response.body()) def main(): # 使用with上下文管理,自动处理playwright的启动/停止 with sync_playwright() as playwright: browser = playwright.chromium.launch(headless=True) context = browser.new_context(no_viewport=True) page = context.new_page() # 直接绑定响应处理函数,无需lambda page.on('response', handle_response) page.goto("https://www.booking.com/hotel/it/hotelnordroma.en-gb.html?aid=304142&checkin=2025-05-15&checkout=2025-05-16#map_opened-map_trigger_header_pin") page.wait_for_timeout(1000) browser.close() if __name__=='__main__': main()
关键修改点
- 用
with sync_playwright()替代手动start()/stop(),代码更简洁且自动释放资源。 - 新增响应类型判断,仅处理文本类资源,避免二进制解码报错。
- 将文件打开模式从
w(覆盖)改为a(追加),保留所有文本响应内容。 - 可选添加二进制资源的保存逻辑(示例中注释了图片保存代码)。
内容的提问来源于stack exchange,提问作者Mojsa
相关产品推荐
相关产品推荐

