HXT解析含可选BOM的HTTP响应XML:如何避免报错?
问题描述
使用HXT解析http-conduit请求返回的XML响应体时,因部分API版本的响应体开头带有UTF-8字节顺序标记(BOM),触发解析错误:
error: ""\65279<?xml version="1.0" encoding="utf-8"?><Enume..."" (line 1, column 1): unexpected "\65279" expecting xml declaration, comment, processing instruction, "<!DOCTYPE" or "<"
现有方案是先解析原响应体,失败则去掉首字符重试,但存在BOM存在时会输出报错信息的问题:
... let resBody = Data.ByteString.UTF8.toString . toStrict $ getResponseBody response parseBody body = runX $ readString [withValidate no] body >>> getChildren >>> ... xs <- parseBody resBody val <- case xs of x : _ -> pure x _ -> head <$> (parseBody $ drop 1 resBody) ...
无报错解析方案
方案一:主动移除UTF-8 BOM(推荐)
UTF-8的BOM是固定字节序列0xEFBBBF,可以在转换为字符串前直接检查并移除,从根源避免解析错误:
import Data.ByteString (ByteString) import qualified Data.ByteString as BS -- 定义UTF-8 BOM的字节序列 utf8Bom :: ByteString utf8Bom = BS.pack [0xEF, 0xBB, 0xBF] -- 移除UTF-8 BOM的工具函数 stripUtf8Bom :: ByteString -> ByteString stripUtf8Bom bs = if BS.isPrefixOf utf8Bom bs then BS.drop 3 bs else bs -- 处理响应体并解析 let strictBody = toStrict $ getResponseBody response cleanedBody = Data.ByteString.UTF8.toString $ stripUtf8Bom strictBody val <- runX $ readString [withValidate no] cleanedBody >>> getChildren >>> ...
该方案无需重试逻辑,也不会触发任何解析错误输出。
方案二:利用HXT编码选项自动处理BOM
HXT的readString支持通过编码配置自动识别并跳过BOM,指定withEncoding utf8即可:
import Text.XML.HXT.Core (withEncoding, utf8) val <- runX $ readString [withValidate no, withEncoding utf8] (Data.ByteString.UTF8.toString . toStrict $ getResponseBody response) >>> getChildren >>> ...
HXT的编码模块会自动检测UTF-8 BOM并忽略,无需额外手动处理。
方案三:捕获解析错误静默处理(不推荐)
如果要保留重试逻辑,可以通过捕获异常避免错误信息输出,但本质还是会触发一次解析错误,仅做静默处理:
import Control.Exception (catch, SomeException) let resBody = Data.ByteString.UTF8.toString . toStrict $ getResponseBody response parseBody body = runX $ readString [withValidate no] body >>> getChildren >>> ... val <- catch (head <$> parseBody resBody) (\(_::SomeException) -> head <$> parseBody (drop 1 resBody))
内容的提问来源于stack exchange,提问作者Sledge
相关产品推荐
相关产品推荐

