You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python Playwright爬虫遇UnicodeDecodeError:如何识别响应类型并处理?

Playwright爬虫:响应类型判断及UnicodeDecodeError解决

问题场景

刚上手Python+Playwright开发网页爬虫,运行代码时触发UnicodeDecodeError,需要明确如何判断响应数据类型(文本/二进制等),并解决将响应转文本保存的报错问题。

原代码

from playwright.sync_api import sync_playwright
import json

def handle_response(response):   
  with open("copy.txt", "w", encoding="utf-8") as file:
      file.write(response.text()) 

def main():
  playwright=sync_playwright().start()
  browser=playwright.chromium.launch(headless=True)
  browser.new_context(no_viewport=True)
  page=browser.new_page()  
  page.on('response',lambda response:handle_response(response))  
  page.goto("https://www.booking.com/hotel/it/hotelnordroma.en-gb.html?aid=304142&checkin=2025-05-15&checkout=2025-05-16#map_opened-map_trigger_header_pin")    
  page.wait_for_timeout(1000)   
  browser.close()
  playwright.stop()

if __name__=='__main__':
   main()

报错信息

Exception has occurred: UnicodeDecodeError
'utf-8' codec can't decode byte 0x89 in position 0: invalid start byte
File "J:\SeSa\Playwright\sample.py", line 6, in handle_response
file.write(response.text())

File "J:\SeSa\Playwright\sample.py", line 15, in 
page.on('response',lambda response:handle_response(response))<br />
~~~~~~~~~~~~~~~~^^^^^^^^^^
UnicodeDecodeError: 'utf-8' codec can't decode byte 0x89 in position 0: invalid start byte

问题分析

报错里的0x89是PNG图片的起始字节,说明代码尝试将二进制响应(比如图片、字体、视频等非文本资源)强制解码为UTF-8文本,自然会触发解码失败。


解决方案

1. 如何判断响应数据类型

  • 通过response.content_type属性:直接获取响应的MIME类型,比如text/html(网页)、application/json(接口数据)属于文本类型;image/png(图片)、application/octet-stream(通用二进制)属于二进制类型。
  • 通过响应头Content-Type:和上面等价,可通过response.headers.get("Content-Type")获取。
  • 方法区分:response.text()会自动按响应编码解码为字符串,仅适用于文本资源;response.body()返回原始二进制字节,适用于所有类型的响应。

2. 修改代码解决报错

核心逻辑:只处理文本类型的响应,跳过或单独处理二进制资源。以下是优化后的代码:

from playwright.sync_api import sync_playwright

def handle_response(response):
    content_type = response.content_type
    # 筛选需要处理的文本类型响应(可根据需求调整类型范围)
    text_types = ("text/", "application/json", "application/javascript")
    if any(content_type.startswith(t) for t in text_types):
        try:
            text_content = response.text()
            # 用追加模式写入,避免覆盖之前的内容
            with open("copy.txt", "a", encoding="utf-8") as file:
                file.write(f"=== 响应URL: {response.url} ===\n")
                file.write(text_content + "\n\n")
        except Exception as e:
            print(f"处理文本响应失败[{response.url}]: {str(e)}")
    else:
        # 二进制资源可选择跳过,或保存为二进制文件
        print(f"跳过二进制响应[{response.url}]: {content_type}")
        # 示例:保存图片类二进制资源
        # if content_type.startswith("image/"):
        #     filename = response.url.split("/")[-1]
        #     with open(filename, "wb") as img_file:
        #         img_file.write(response.body())

def main():
    # 使用with上下文管理,自动处理playwright的启动/停止
    with sync_playwright() as playwright:
        browser = playwright.chromium.launch(headless=True)
        context = browser.new_context(no_viewport=True)
        page = context.new_page()
        # 直接绑定响应处理函数,无需lambda
        page.on('response', handle_response)
        page.goto("https://www.booking.com/hotel/it/hotelnordroma.en-gb.html?aid=304142&checkin=2025-05-15&checkout=2025-05-16#map_opened-map_trigger_header_pin")
        page.wait_for_timeout(1000)
        browser.close()

if __name__=='__main__':
    main()

关键修改点

  • 用with sync_playwright()替代手动start()/stop(),代码更简洁且自动释放资源。
  • 新增响应类型判断,仅处理文本类资源,避免二进制解码报错。
  • 将文件打开模式从w(覆盖)改为a(追加),保留所有文本响应内容。
  • 可选添加二进制资源的保存逻辑(示例中注释了图片保存代码)。

内容的提问来源于stack exchange,提问作者Mojsa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.14 01:15:00