使用PHP解析opendata.ndw.nu的xml.gz文件时返回空对象,求指引
解决PHP解析xml.gz后SimpleXMLElement为空的问题
可能的原因与解决方向
1. 处理XML命名空间
多数公共开放数据的XML会包含命名空间,SimpleXML默认不会自动解析命名空间下的节点,导致print_r显示空对象,但实际数据存在。
测试并处理命名空间:
<?php $content = file_get_contents("http://opendata.ndw.nu/brugopeningen.xml.gz"); $uncompressed = gzdecode($content); libxml_use_internal_errors(true); $xml = simplexml_load_string($uncompressed); // 获取所有命名空间 $namespaces = $xml->getNamespaces(true); print_r($namespaces); // 若存在命名空间,用children方法访问对应节点 if (!empty($namespaces)) { // 取第一个命名空间(根据实际XML调整) $targetNs = reset($namespaces); $dataNodes = $xml->children($targetNs); print_r($dataNodes); } ?>
2. 捕获XML解析错误
开启libxml错误捕获,查看是否有格式或编码问题:
<?php $content = file_get_contents("http://opendata.ndw.nu/brugopeningen.xml.gz"); $uncompressed = gzdecode($content); libxml_use_internal_errors(true); $xml = simplexml_load_string($uncompressed); if (!$xml) { foreach (libxml_get_errors() as $error) { echo "解析错误:" . $error->message . " 行号:" . $error->line . "\n"; } libxml_clear_errors(); } ?>
常见错误包括编码不匹配(比如XML声明是UTF-16但PHP按UTF-8解析)、标签未闭合等。
3. 改用流式解压避免内存问题
如果文件较大,file_get_contents可能导致内存溢出,改用流式读取解压:
<?php $handle = gzopen('http://opendata.ndw.nu/brugopeningen.xml.gz', 'rb'); $xmlContent = ''; while (!gzeof($handle)) { $xmlContent .= gzread($handle, 4096); } gzclose($handle); libxml_use_internal_errors(true); $xml = simplexml_load_string($xmlContent); // 后续处理同前 ?>
4. 用DOMDocument辅助解析
若SimpleXML无法处理,可先用DOMDocument加载再转成SimpleXML:
<?php $content = file_get_contents("http://opendata.ndw.nu/brugopeningen.xml.gz"); $uncompressed = gzdecode($content); $dom = new DOMDocument(); $dom->loadXML($uncompressed); $xml = simplexml_import_dom($dom); print_r($xml); ?>
内容的提问来源于stack exchange,提问作者JesperV00
相关产品推荐
相关产品推荐

