如何让PHP爬虫等待JS加载完成后再抓取第三方站点的动态OG标签
问题修复方案
你当前使用的phpscraper是静态爬虫,仅能抓取页面初始HTML内容,无法执行JavaScript渲染,因此拿不到动态生成的OG标签。修复需要更换为支持无头浏览器渲染的爬虫方案,具体改造步骤如下:
改造前提
服务器需要提前安装Chromium/Chrome浏览器,用于无头渲染执行页面JS。
步骤1:替换爬虫依赖
推荐使用封装更简单的drupal/puppeteer库,通过composer安装:
composer require drupal/puppeteer
步骤2:修改核心抓取逻辑
原有代码还有一处逻辑错误:$description = $web->openGraph['og:title'];赋值对象错误,应改为给$title赋值,修改后的完整代码如下:
<?php define( 'DELAY', '0' ); // delay to redirect define('OG_CACHE_TTL', 3600); // 增加OG数据缓存时间,单位秒,避免重复渲染浪费资源 yourls_add_action( 'pre_redirect', 'og_scraper' ); function og_scraper( $args ) { $url = $args[0]; $parsed_url = parse_url($url); $cache_key = 'og_' . md5($url); // 优先读缓存,提升性能 $cached_data = apc_fetch($cache_key); if ($cached_data) { $data = $cached_data; } else { require '/var/www/fantasticpix.com/htdocs/pixiviz/dist/artworks/includes/vendor/autoload.php'; // 启动无头浏览器 $puppeteer = new \Drupal\puppeteer\Puppeteer(); $options = [ 'waitUntil' => 'networkidle2', // 等待网络空闲再读取内容,确保JS执行完成 'timeout' => 10000, 'args' => ['--no-sandbox', '--disable-setuid-sandbox'] // Linux环境下权限兼容参数 ]; // 渲染页面 $page = $puppeteer->setOptions($options)->goto($url); // 读取页面渲染后的完整HTML $html = $page->getContent(); // 解析OG标签 $doc = new \DOMDocument(); @$doc->loadHTML($html); $xpath = new \DOMXPath($doc); $data = []; $metas = $xpath->query('//meta[@property]'); foreach ($metas as $meta) { $property = $meta->getAttribute('property'); if (str_starts_with($property, 'og:')) { $data[$property] = $meta->getAttribute('content'); } } // 兜底取常规meta和title if (empty($data['og:description'])) { $desc_nodes = $xpath->query('//meta[@name="description"]'); if ($desc_nodes->length > 0) { $data['og:description'] = $desc_nodes->item(0)->getAttribute('content'); } } if (empty($data['og:title'])) { $title_nodes = $xpath->query('//title'); if ($title_nodes->length > 0) { $data['og:title'] = $title_nodes->item(0)->nodeValue; } } // 处理图片相对路径 if (!empty($data['og:image']) && !str_starts_with($data['og:image'], 'http')) { $data['og:image'] = $parsed_url['scheme'] . "://" . $parsed_url['host'] . $data['og:image']; } // 写入缓存 apc_store($cache_key, $data, OG_CACHE_TTL); } $title = $data['og:title'] ?? ''; $description = $data['og:description'] ?? ''; $image = $data['og:image'] ?? ''; ?> <html> <head> <?php echo '<meta property="og:title" content="' . htmlspecialchars($title) . '" /> <meta property="og:type" content="' . htmlspecialchars($data['og:type'] ?? 'website') . '" /> <meta property="og:url" content="' . htmlspecialchars($data['og:url'] ?? $url) . '" /> <meta property="og:image" content="' . htmlspecialchars($image) . '" /> <meta property="og:site_name" content="' . htmlspecialchars($data['og:site_name'] ?? '') . '" /> <meta property="og:description" content="' . htmlspecialchars($description) . '" />'; ?> </head> <body><iframe id="yourls-frame" frameborder="0" noresize="noresize" src="<?php echo htmlspecialchars($url); ?>" name="yourlsFrame" style="position:fixed; top:0; left:0; bottom:0; right:0; width:100%; height:100%; border:none; margin:0; padding:0; overflow:hidden; z-index:999999;"></iframe></body> </html> <?php die(); }
注意事项
- 示例中使用了PHP8+的
str_starts_with函数,如果你使用PHP7及以下版本,将其替换为substr($str, 0, strlen($prefix)) === $prefix的判断逻辑即可 - 示例中使用了apc做缓存,你也可以换成Redis、文件缓存等其他缓存方案,避免每次请求都启动无头浏览器,大幅提升响应速度
- 如果服务器内存较小,可以调低无头浏览器的并发数量,避免资源耗尽
- 输出OG内容前用
htmlspecialchars转义,避免XSS漏洞
内容的提问来源于stack exchange,提问作者TravelWhere
相关产品推荐
相关产品推荐

