使用PHP simple_html_dom解析HTML时如何移除元素内span标签仅提取价格
方案1:DOM操作移除span标签(推荐,兼容性更好)
直接在提取文本前删除.price节点内的单位span元素,避免匹配错误,实现代码如下:
require('simple_html_dom.php'); $html = file_get_html("www.example.com"); // 提取柴油价格 $dieselPriceNode = $html->find(".on .price", 0); if ($unitSpan = $dieselPriceNode->find("span", 0)) { $unitSpan->remove(); } $info['diesel'] = trim($dieselPriceNode->innertext); // 提取95号汽油价格 $pb95PriceNode = $html->find(".pb .price", 0); if ($unitSpan = $pb95PriceNode->find("span", 0)) { $unitSpan->remove(); } $info['pb95'] = trim($pb95PriceNode->innertext); // 提取液化石油气价格 $lpgPriceNode = $html->find(".lpg .price", 0); if ($unitSpan = $lpgPriceNode->find("span", 0)) { $unitSpan->remove(); } $info['lpg'] = trim($lpgPriceNode->innertext);
方案2:正则匹配提取价格(代码更简洁)
如果确认页面结构不会变动,也可以直接从原始文本中匹配数字+逗号组成的价格部分:
require('simple_html_dom.php'); $html = file_get_html("www.example.com"); // 匹配规则:提取连续的数字和逗号组合 preg_match('/[\d,]+/', $html->find(".on .price", 0)->innertext, $match); $info['diesel'] = $match[0] ?? ''; preg_match('/[\d,]+/', $html->find(".pb .price", 0)->innertext, $match); $info['pb95'] = $match[0] ?? ''; preg_match('/[\d,]+/', $html->find(".lpg .price", 0)->innertext, $match); $info['lpg'] = $match[0] ?? '';
内容的提问来源于stack exchange,提问作者Zimnytu
相关产品推荐
相关产品推荐

