使用PHP Simple HTML Dom Parser无法抓取网页预测数据的问题
问题分析与解决方案
核心问题
_ngcontent-coincodex-c2362887433是Angular框架生成的动态属性,每次页面部署、刷新后都可能随机变化,用它作为选择器完全不可靠,这是你无法获取预期内容的根本原因。
解决方案:改用稳定选择器
目标元素有固定的class prediction-range,可以通过这个稳定标识定位父容器,再提取内部的预测数据:
require_once 'simple_html_dom.php'; $url = 'https://coincodex.com/crypto/bad-idea-ai/price-prediction'; $html = file_get_html($url); // 遍历所有预测模块容器 foreach($html->find('div.prediction-range') as $container) { // 获取预测周期(如5-Day Prediction) $period = trim($container->find('div', 0)->plaintext); // 获取包含数值的文本并清理空格 $rawValue = trim($container->find('div', 1)->plaintext); // 提取纯数值部分(过滤箭头符号) preg_match('/\$ ([\d.₇]+)/', $rawValue, $matches); $predictionValue = isset($matches[1]) ? $matches[1] : '无法获取数值'; echo "{$period}: {$predictionValue}\n"; }
额外说明
如果运行后仍无法获取数据,大概率是页面内容通过JavaScript动态渲染的——file_get_html只能抓取原始静态HTML,无法获取JS加载后的内容。这种情况下需要使用支持JS执行的工具(如PhantomJS、PHP版Puppeteer)来抓取渲染后的页面。
内容的提问来源于stack exchange,提问作者Blackstone
相关产品推荐
相关产品推荐

