You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在XML对比中忽略文本具体值仅校验数据类型?

XML对比:仅校验文本数据类型而非具体值

问题场景

尝试使用XMLUnit的CompareMatcher进行XML对比,代码如下:

import org.junit.jupiter.api.Test;
import org.xmlunit.diff.DefaultNodeMatcher;
import org.xmlunit.diff.ElementSelectors;
import org.xmlunit.matchers.CompareMatcher;
import static org.hamcrest.MatcherAssert.assertThat;

public class XmlDemo4 {

    @Test
    public void demoMethod() {
        String actual = "<struct><int>3</int><boolean>false</boolean></struct>";
        String expected = "<struct><boolean>false</boolean><int>4</int></struct>";

        assertThat(actual, CompareMatcher.isSimilarTo(expected)
                .ignoreWhitespace().normalizeWhitespace().
                withNodeMatcher(new 
        DefaultNodeMatcher(ElementSelectors.byName,ElementSelectors.Default)));

    }

}

运行后因文本值不匹配报错:

java.lang.AssertionError: 
Expected: Expected text value '4' but was '3' - comparing <int ...>4</int> at 
/struct[1]/int[1]/text()[1] to <int ...>3</int> at /struct[1]/int[1]/text()[1]:
<int>4</int>
     but: result was: 
<int>3</int>

at org.hamcrest.MatcherAssert.assertThat(MatcherAssert.java:20)
at org.hamcrest.MatcherAssert.assertThat(MatcherAssert.java:8)
at StringXml.XmlDemo4.demoMethod(XmlDemo4.java:29)
at java.base/java.util.ArrayList.forEach(ArrayList.java:1541)
at java.base/java.util.ArrayList.forEach(ArrayList.java:1541)

Process finished with exit code -1

需求:实现仅对比文本内容的数据类型,不校验具体值的XML对比。

实现方案

XMLUnit允许通过自定义DifferenceEvaluator覆盖默认对比逻辑,我们可以在其中添加类型校验规则,忽略值的差异只要两边数据类型匹配。

修改后的代码

import org.junit.jupiter.api.Test;
import org.xmlunit.diff.*;
import org.xmlunit.matchers.CompareMatcher;
import static org.hamcrest.MatcherAssert.assertThat;

public class XmlDemo4 {

    @Test
    public void demoMethod() {
        String actual = "<struct><int>3</int><boolean>false</boolean></struct>";
        String expected = "<struct><boolean>false</boolean><int>4</int></struct>";

        assertThat(actual, CompareMatcher.isSimilarTo(expected)
                .ignoreWhitespace()
                .normalizeWhitespace()
                .withNodeMatcher(new DefaultNodeMatcher(ElementSelectors.byName))
                .withDifferenceEvaluator((comparison, outcome) -> {
                    // 仅处理文本值不相等的情况
                    if (outcome != ComparisonResult.EQUAL && comparison.getType() == ComparisonType.TEXT_VALUE) {
                        // 从XPath中提取节点名称
                        String elementName = comparison.getControlDetails().getXPath()
                                .split("/")[2].split("\\[")[0];
                        String controlText = comparison.getControlDetails().getValue().toString();
                        String testText = comparison.getTestDetails().getValue().toString();

                        // 根据节点标签校验数据类型
                        switch (elementName) {
                            case "int":
                                try {
                                    Integer.parseInt(controlText);
                                    Integer.parseInt(testText);
                                    // 两边都是整数,忽略值差异
                                    return ComparisonResult.EQUAL;
                                } catch (NumberFormatException e) {
                                    // 类型不匹配,保留差异
                                    return outcome;
                                }
                            case "boolean":
                                boolean controlIsBool = isBoolean(controlText);
                                boolean testIsBool = isBoolean(testText);
                                if (controlIsBool && testIsBool) {
                                    return ComparisonResult.EQUAL;
                                }
                                break;
                            // 可扩展支持更多类型:float、date等
                        }
                    }
                    // 其他情况保持原对比结果
                    return outcome;
                }));
    }

    // 辅助方法:判断字符串是否为布尔值(支持大小写)
    private boolean isBoolean(String text) {
        return "true".equalsIgnoreCase(text) || "false".equalsIgnoreCase(text);
    }
}

逻辑说明

  1. 拦截文本差异:只处理ComparisonType.TEXT_VALUE类型的不相等结果,其他差异(如节点顺序、属性等)保持原逻辑。
  2. 提取节点名称:从对比结果的XPath中解析出当前节点的标签名(如int、boolean)。
  3. 类型校验:
    • 对<int>节点:尝试将两边文本转为整数,成功则认为类型匹配,忽略值差异。
    • 对<boolean>节点:判断两边文本是否为合法布尔值(支持大小写),是则忽略值差异。
  4. 扩展支持:可根据需求添加更多类型的校验逻辑(如浮点数、日期格式等)。

内容的提问来源于stack exchange,提问作者KB1

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 04:45:43