You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

XML转PHP数组后丢失属性数据,基于指定GitHub函数的问题求助

解决XML转PHP数组时属性丢失的问题

我之前也踩过这个坑——你这个问题的核心原因就是domNodeToArray方法没处理DOM节点的属性(Attributes),默认的DOM节点遍历只会处理子节点和文本,不会主动提取属性。下面是修复后的完整实现,我会把关键改动点给你讲清楚:

修复后的完整代码

private function parseXml($xmlString) {
    $doc = new \DOMDocument;
    // 可选:禁用XML格式错误提示,避免小格式问题中断执行
    $doc->loadXML($xmlString, LIBXML_NOERROR | LIBXML_NOWARNING);
    $root = $doc->documentElement;
    $output[$root->tagName] = $this->domNodeToArray($root);
    return $output;
}

/**
 * @param \DOMNode $node
 * @return array|string
 */
private function domNodeToArray(\DOMNode $node) {
    $output = [];

    // 第一步:提取节点的所有属性
    if ($node->hasAttributes()) {
        $attributes = [];
        foreach ($node->attributes as $attr) {
            $attributes[$attr->nodeName] = $attr->nodeValue;
        }
        // 用@attributes作为属性键,是XML转数组的通用约定,方便区分属性和节点内容
        $output['@attributes'] = $attributes;
    }

    // 第二步:处理子节点和文本内容
    if ($node->hasChildNodes()) {
        $children = [];
        
        foreach ($node->childNodes as $child) {
            // 忽略空白文本节点(比如XML里的换行、空格),避免干扰结果
            if ($child instanceof \DOMText && trim($child->nodeValue) === '') {
                continue;
            }
            
            // 递归处理元素节点
            if ($child instanceof \DOMElement) {
                $childData = $this->domNodeToArray($child);
                // 处理重复标签:如果同一父节点下有多个同名子节点,自动转为索引数组
                if (isset($children[$child->tagName])) {
                    if (!is_array($children[$child->tagName]) || !isset($children[$child->tagName][0])) {
                        $children[$child->tagName] = [$children[$child->tagName]];
                    }
                    $children[$child->tagName][] = $childData;
                } else {
                    $children[$child->tagName] = $childData;
                }
            } 
            // 处理非空白的文本节点
            elseif ($child instanceof \DOMText) {
                $children['@value'] = trim($child->nodeValue);
            }
        }

        // 合并属性和子节点/文本内容
        $output = array_merge($output, $children);
        
        // 优化:如果只有纯文本内容,直接返回字符串(避免多余的数组嵌套)
        if (count($output) === 1 && isset($output['@value'])) {
            return $output['@value'];
        }
    } 
    // 没有子节点但有文本内容的情况
    else {
        $textContent = trim($node->textContent);
        if ($textContent !== '') {
            if (!empty($output)) {
                // 有属性的情况下,文本存入@value
                $output['@value'] = $textContent;
            } else {
                // 无属性的纯文本节点,直接返回字符串
                return $textContent;
            }
        }
    }

    return $output;
}

关键改动说明

  1. 属性提取逻辑:通过$node->hasAttributes()判断节点是否有属性,遍历所有属性存入@attributes数组——这是XML转数组的通用约定,能清晰区分属性和节点内容。
  2. 空白节点过滤:忽略XML中的换行、空格等空白文本节点,避免生成多余的无效数据。
  3. 重复标签处理:如果同一父节点下有多个同名子节点(比如多个<item>),自动转为索引数组,不会出现后面的节点覆盖前面的情况。
  4. 文本节点优化:区分纯文本节点和带属性的文本节点,减少不必要的数组嵌套,让结果更简洁。

示例效果

比如输入以下XML:

<book id="123" category="fiction">
    <title>Harry Potter</title>
    <author>J.K. Rowling</author>
    <price currency="USD">29.99</price>
</book>

转成PHP数组后会得到:

[
    'book' => [
        '@attributes' => [
            'id' => '123',
            'category' => 'fiction'
        ],
        'title' => 'Harry Potter',
        'author' => 'J.K. Rowling',
        'price' => [
            '@attributes' => ['currency' => 'USD'],
            '@value' => '29.99'
        ]
    ]
]

你可以根据自己的需求调整键名(比如把@attributes改成attributes,@value改成value),或者修改空白节点的处理逻辑。

内容的提问来源于stack exchange,提问作者aborted

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:04:17