You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

NiFi使用Python ExecuteScript转JSON为XML时非英文字符乱码如何解决

修复方案

问题根源

乱码是由三处编码处理错误共同导致的:

  • 给Att4的text属性赋值时额外执行了encode("utf8")操作,将字符串转为字节对象,ElementTree会直接将字节的十进制值转换为XML实体,最终显示为乱码
  • ET.tostring()默认使用ASCII编码输出,所有非ASCII字符都会被自动转义为实体编码
  • 输出时额外做了bytearray转换,可能引发额外的编码冲突

修正后的完整代码

import json
import xml.etree.ElementTree as ET
import java.io
from org.apache.commons.io import IOUtils
from java.nio.charset import StandardCharsets
from org.apache.nifi.processor.io import StreamCallback

class ModJSON(StreamCallback):

    def __init__(self):
        pass

    def process(self, inputStream, outputStream):
        text = IOUtils.toString(inputStream, StandardCharsets.UTF_8)
        data = json.loads(text)
        root = ET.Element("headerinfo")
        entity = ET.SubElement(root, "headerfile")
        ET.SubElement(entity, "Att1").text = str(data["Header"]["Att1"])
        ET.SubElement(entity, "Att2").text = str(data["Header"]["Att2"])
        ET.SubElement(entity, "Att3").text = str(data["Header"]["Att3"])
        # 直接赋值字符串,不要提前编码
        ET.SubElement(entity, "Att4").text = data["Header"]["Att4"]
        # 指定UTF-8编码,同时添加XML声明标明编码
        xmlNew = ET.tostring(root, encoding='utf-8', xml_declaration=True)
        # 直接输出编码后的字节流
        outputStream.write(xmlNew)

flowFile = session.get()
if flowFile != None:
    try :
        flowFile = session.write(flowFile, ModJSON())
        flowFile = session.putAttribute(flowFile, "filename", 'headerfile.xml')
        session.transfer(flowFile, REL_SUCCESS)
        session.commit()
    except Exception as e:
        flowFile = session.putAttribute(flowFile,'python_error', str(e))
        session.transfer(flowFile, REL_FAILURE)

补充说明

如果后续在NiFi下游处理器、或者本地打开生成的XML文件仍有乱码,需要确认对应工具的默认字符集配置为UTF-8即可。

内容的提问来源于stack exchange,提问作者maverick639

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.01 07:18:03