You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

FastAPI返回XML含字符:如何去除并正常解析?

解决UTF-8 BOM()问题并正确处理XML响应

原因分析

是UTF-8编码的字节顺序标记(BOM),当外部API返回带BOM的UTF-8内容时,requests默认的编码检测未识别utf-8-sig格式,导致BOM被当作普通字符保留在字符串中。


方案1:让requests自动去除BOM

直接给响应指定utf-8-sig编码,response.text会自动剥离BOM:

def __post():
    headers = {'Content-Type': 'application/xml'}
    auth = requests.auth.HTTPBasicAuth("username", "password")
  
    response = requests.post('https://example.org/id', headers=headers, auth=auth)
    # 指定编码为utf-8-sig,自动处理BOM
    response.encoding = 'utf-8-sig'
    return response.text

方案2:手动去除字符串中的BOM

如果已经拿到带BOM的字符串,可通过检查并剥离开头的BOM字符:

def strip_utf8_bom(xml_str: str) -> str:
    utf8_bom = '\ufeff'
    return xml_str.lstrip(utf8_bom) if xml_str.startswith(utf8_bom) else xml_str

# 在接口中使用
@router.get("/{id}")
def test():
    xml = __post()
    xml_clean = strip_utf8_bom(xml)
    return Response(content=xml_clean, media_type="application/xml")

保存文件的正确方式

已去除BOM后,保存文件时使用utf-8编码即可(无需再用utf-8-sig,否则会重新添加BOM):

with open('response.xml', 'w', encoding='utf-8') as f:
     f.write(xml_clean)

解析XML响应

处理完BOM后,直接用Python标准库或第三方库解析即可:

import xml.etree.ElementTree as ET

# 解析清理后的XML字符串
root = ET.fromstring(xml_clean)

# 或者从文件解析
tree = ET.parse('response.xml')
root = tree.getroot()

内容的提问来源于stack exchange,提问作者oliverbj

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.23 03:45:42