You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用file_get_contents获取class为c-shadow的链接的所有锚文本

用PHP获取指定class链接的锚文本

嘿,这个需求用PHP的DOM扩展来实现是最稳妥的——别用正则处理HTML,坑太多!我给你整理了一套完整的实现方案,分步骤来看:

核心思路

先用file_get_contents抓取目标页面的HTML源码,再通过DOMDocument和DOMXPath解析页面,精准定位所有class为c-shadow的<a>标签,最后提取它们的锚文本。

完整代码示例

<?php
// 目标页面URL
$targetUrl = 'https://example.com/your-target-page';

// 1. 抓取页面内容(带基础错误处理)
$htmlContent = file_get_contents($targetUrl);
if (!$htmlContent) {
    die('无法获取目标页面内容,请检查URL是否正确或服务器是否允许访问');
}

// 2. 初始化DOM解析器,处理不规范HTML的警告
libxml_use_internal_errors(true); // 屏蔽HTML解析时的警告
$dom = new DOMDocument();
$dom->loadHTML($htmlContent);
libxml_clear_errors(); // 清除错误缓存

// 3. 用XPath定位目标链接
$xpath = new DOMXPath($dom);
// 这里的XPath写法确保精准匹配class包含"c-shadow"的a标签(避免匹配类似c-shadow-xxx的class)
$links = $xpath->query('//a[contains(concat(" ", normalize-space(@class), " "), " c-shadow ")]');

// 4. 提取锚文本并收集结果
$anchorTexts = [];
if ($links->length > 0) {
    foreach ($links as $link) {
        // 获取纯文本内容,自动去除多余空格和换行
        $anchorText = trim($link->textContent);
        if (!empty($anchorText)) {
            $anchorTexts[] = $anchorText;
        }
    }
}

// 输出结果(或者根据你的需求处理)
print_r($anchorTexts);
?>

注意事项

  • 关于file_get_contents的局限性:有些服务器会阻止file_get_contents的请求(因为默认的User-Agent太直白),这时候可以改用CURL来模拟浏览器请求,比如设置User-Agent头:
    $ch = curl_init();
    curl_setopt($ch, CURLOPT_URL, $targetUrl);
    curl_setopt($ch, CURLOPT_RETURNTRANSFER, true);
    curl_setopt($ch, CURLOPT_USERAGENT, 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36');
    $htmlContent = curl_exec($ch);
    curl_close($ch);
    
  • XPath的精准匹配:为什么要用concat(" ", normalize-space(@class), " ")?因为class属性可能有多个值(比如class="c-shadow other-class"),这样写能确保我们匹配的是独立的c-shadow类,而不是其他包含这个字符串的类名。

内容的提问来源于stack exchange,提问作者wrkman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 08:11:16