OpenSAML序列化UTF字符时代理对被转成双NCR的问题排查
问题:XML序列化中Unicode补充平面字符的编码异常
负载中包含类似𠮷的Unicode补充平面字符,将其序列化为XML时,预期结果为以下两种之一:
- 不编码为数值字符引用(NCR),直接输出原字符
- 编码为单个NCR形式:
𠮷
但实际观察到的是该字符被拆分为两个代理对的NCR:��。
复现代码
package something...; import net.shibboleth.utilities.java.support.xml.SerializeSupport; import org.opensaml.core.config.InitializationException; import org.opensaml.core.config.InitializationService; import org.opensaml.core.xml.schema.XSString; import org.opensaml.core.xml.schema.impl.XSStringBuilder; import org.opensaml.core.xml.util.XMLObjectSupport; import org.opensaml.saml.saml2.core.Attribute; import org.opensaml.saml.saml2.core.AttributeValue; import org.w3c.dom.Document; /** Generates Response.xml using openSAML library version 3. */ final class SamlResponseGenerator { public static void main(String[] args) throws Exception { try { InitializationService.initialize(); } catch (InitializationException e) { throw new IllegalStateException(e); } // First Name Attribute. Attribute firstNameAttribute = (Attribute) XMLObjectSupport.buildXMLObject(Attribute.DEFAULT_ELEMENT_NAME); firstNameAttribute.setName("FirstName"); XSStringBuilder firstNameTestValueStringBuilder = new XSStringBuilder(); XSString firstNameAttributeValueXS = firstNameTestValueStringBuilder.buildObject( AttributeValue.DEFAULT_ELEMENT_NAME, XSString.TYPE_NAME); String myName = "𠮷"; firstNameAttributeValueXS.setValue(myName); firstNameAttribute.getAttributeValues().add(firstNameAttributeValueXS); Document doc = XMLObjectSupport.marshall(firstNameAttribute).getOwnerDocument(); String docString = SerializeSupport.nodeToString(doc); System.out.println(docString); } private SamlResponseGenerator() {} }
已尝试的解决方案
- 使用Transformer序列化(无效)
- 结合UTF-8编码的FileWriter与XMLHelper.writeNode(无效)
- 对XML输出进行后处理(有效,但属于临时方案)
- 测试OpenSAML v2与v3版本(均无效)
- 查阅代理对XML序列化相关技术资料
疑问
- 是否因使用错误的序列化器或配置导致该问题?如何配置才能得到符合预期的输出(与旧服务器行为一致)?
- 维基百科提到代理对不允许出现在数值字符引用表示法中,这一点是否与当前问题直接相关?
内容的提问来源于stack exchange,提问作者Mohit Jain
相关产品推荐
相关产品推荐

