You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PHP实现网页URL与锚文本提取:DOM及cURL方案求助

Hey there! Let's work through your PHP web scraping needs step by step. I'll start with fixing your DOMDocument code (solving the blank page and warning issues) to output both URLs and anchor text, then move on to the cURL-only solutions you asked for—both regex and non-regex versions.

Fixed DOMDocument Solution (Fix Warnings + Output URL & Anchor Text)

Your blank page likely came from unhandled cURL failures (e.g., blocked requests, network issues) or invalid HTML breaking the parser. The @ suppression for warnings is also not ideal—we'll use proper libxml error handling instead. Here's the improved code:

<?php
// Target URL (replace with your desired page)
$url = "http://www.example.com";
$ch = curl_init();
$timeout = 10;

// Configure cURL to avoid being blocked and handle errors
curl_setopt($ch, CURLOPT_URL, $url);
curl_setopt($ch, CURLOPT_RETURNTRANSFER, 1);
curl_setopt($ch, CURLOPT_CONNECTTIMEOUT, $timeout);
// Add a user-agent to mimic a real browser (critical for many sites)
curl_setopt($ch, CURLOPT_USERAGENT, "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36");

$html = curl_exec($ch);

// Check if cURL request succeeded (prevents blank pages)
if ($html === false) {
    die("cURL Error: " . curl_error($ch));
}
curl_close($ch);

// Handle invalid HTML warnings properly (no more @ suppression)
libxml_use_internal_errors(true);
$dom = new DOMDocument();
$dom->loadHTML($html);
libxml_clear_errors(); // Clear any stored parsing errors

// Output links with their anchor text
echo "<h3>Extracted Links (DOMDocument):</h3>";
foreach ($dom->getElementsByTagName('a') as $link) {
    $href = $link->getAttribute('href');
    // Skip empty or placeholder links
    if (empty($href) || $href === "#") continue;
    
    // Clean anchor text (trim whitespace, handle empty text)
    $anchorText = trim($link->textContent);
    $anchorText = $anchorText ?: $href; // Fallback to URL if no text
    
    echo "<p><strong>URL:</strong> `$href` | <strong>Anchor Text:</strong> $anchorText</p>";
}
?>

Key Fixes & Improvements:

  • cURL Error Checking: Catches failed requests before they break the parser, eliminating blank pages.
  • Proper Warning Handling: Uses libxml_use_internal_errors(true) instead of @ to safely suppress invalid HTML warnings.
  • User-Agent Header: Mimics a real browser to avoid being blocked by sites like Google.
  • Clean Output: Skips invalid links and handles cases where anchor text is missing.
cURL-Only Alternatives (No DOM)

If you want to avoid DOM entirely, here are two solutions using only cURL and string processing:

1. Regex Version (Simple but Limited)

Regex works well for basic HTML structures, but note it can fail with complex, nested HTML. Use this for straightforward pages:

<?php
$url = "http://www.example.com";
$ch = curl_init();
curl_setopt($ch, CURLOPT_URL, $url);
curl_setopt($ch, CURLOPT_RETURNTRANSFER, 1);
curl_setopt($ch, CURLOPT_USERAGENT, "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36");

$html = curl_exec($ch);
if ($html === false) {
    die("cURL Error: " . curl_error($ch));
}
curl_close($ch);

// Regex to match <a> tags (captures href and inner text)
$pattern = '/<a\s+[^>]*href=["\']?([^"\'>]+)["\']?[^>]*>(.*?)<\/a>/is';
preg_match_all($pattern, $html, $matches);

echo "<h3>Extracted Links (Regex):</h3>";
foreach ($matches[1] as $index => $href) {
    $href = trim($href);
    if (empty($href) || $href === "#") continue;
    
    // Clean anchor text (strip nested HTML, trim whitespace)
    $anchorText = trim(strip_tags($matches[2][$index]));
    $anchorText = $anchorText ?: $href;
    
    echo "<p><strong>URL:</strong> `$href` | <strong>Anchor Text:</strong> $anchorText</p>";
}
?>

2. Non-Regex Version (String Manipulation)

This uses raw string functions to parse links—great if you want to avoid regex entirely, but only works for simple HTML:

<?php
$url = "http://www.example.com";
$ch = curl_init();
curl_setopt($ch, CURLOPT_URL, $url);
curl_setopt($ch, CURLOPT_RETURNTRANSFER, 1);
curl_setopt($ch, CURLOPT_USERAGENT, "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36");

$html = curl_exec($ch);
if ($html === false) {
    die("cURL Error: " . curl_error($ch));
}
curl_close($ch);

echo "<h3>Extracted Links (String Manipulation):</h3>";
$pos = 0;
while (($start = strpos($html, '<a ', $pos)) !== false) {
    // Locate the end of the <a> opening tag
    $tagEnd = strpos($html, '>', $start);
    if ($tagEnd === false) break;
    
    // Find href attribute (handles both single and double quotes)
    $hrefStart = strpos($html, 'href="', $start);
    $quote = '"';
    if ($hrefStart === false || $hrefStart > $tagEnd) {
        $hrefStart = strpos($html, "href='", $start);
        $quote = "'";
        if ($hrefStart === false || $hrefStart > $tagEnd) {
            $pos = $tagEnd + 1;
            continue; // Skip <a> tags without href
        }
    }
    
    // Extract the href value
    $hrefStart += 6; // Skip "href=\"" or "href='"
    $hrefEnd = strpos($html, $quote, $hrefStart);
    if ($hrefEnd === false) {
        $pos = $tagEnd + 1;
        continue;
    }
    $href = trim(substr($html, $hrefStart, $hrefEnd - $hrefStart));
    
    // Extract anchor text between opening and closing <a> tags
    $closeTagStart = strpos($html, '</a>', $tagEnd);
    if ($closeTagStart === false) break;
    $anchorText = trim(substr($html, $tagEnd + 1, $closeTagStart - $tagEnd - 1));
    $anchorText = trim(strip_tags($anchorText));
    $anchorText = $anchorText ?: $href;
    
    if (!empty($href) && $href !== "#") {
        echo "<p><strong>URL:</strong> `$href` | <strong>Anchor Text:</strong> $anchorText</p>";
    }
    
    $pos = $closeTagStart + 4; // Move past </a> to avoid reprocessing
}
?>

内容的提问来源于stack exchange,提问作者Php Learner

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 04:19:18