You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用PHP的preg_match_all提取指定HTML内容中的信息?

如何用PHP的preg_match_all提取指定HTML内容?

需求说明

你需要从以下HTML代码中提取每个条目里的4类信息:

  • a标签的href属性值(如/categories/rr/1.html)
  • img标签的src属性值(如http://www.erty.com/images/440f2d2a.jpg)
  • span标签内的文本(如Ind)
  • 括号内的数字(如98)

用户提供的原始HTML代码:

<div class="ti"><div class="pic"> <a href="/categories/rr/1.html"><img src="http://www.erty.com/images/440f2d2a.jpg" alt="Ind"> <span>Ind</span></a> (98) </div></div><div class="ti"><div class="pic"> <a href="/categories/ert/1.html"><img src="http://www.erty.com/images/4123d2b.jpg" alt="Wes"> <span>Wes</span></a> (6044) </div></div>

原正则的问题

你尝试的正则表达式:

preg_match_all('|[^<div class="ti"><div class="pic">].*?[^</div></div>]+|', $test_html, $out, PREG_PATTERN_ORDER);

这个写法逻辑有误:

  • [^<div class="ti"><div class="pic">] 是否定字符集,它匹配的是不是这些单个字符的内容,而不是排除整个标签字符串,这完全不符合你的预期。
  • 整个正则没有针对你需要的4个信息设置捕获组,自然无法提取目标内容。

解决方案

方案1:使用DOMDocument(推荐,处理HTML更可靠)

正则处理HTML容易因为标签格式变化(比如空格、换行)失效,PHP的DOM扩展是更专业的HTML解析工具,示例代码如下:

$html = '<div class="ti"><div class="pic"> <a href="/categories/rr/1.html"><img src="http://www.erty.com/images/440f2d2a.jpg" alt="Ind"> <span>Ind</span></a> (98) </div></div><div class="ti"><div class="pic"> <a href="/categories/ert/1.html"><img src="http://www.erty.com/images/4123d2b.jpg" alt="Wes"> <span>Wes</span></a> (6044) </div></div>';

$dom = new DOMDocument();
// 处理HTML可能存在的格式问题
libxml_use_internal_errors(true);
$dom->loadHTML($html);
libxml_clear_errors();

$xpath = new DOMXPath($dom);
// 定位所有class为ti的div下的pic容器
$picDivs = $xpath->query('//div[@class="ti"]/div[@class="pic"]');

$results = [];
foreach ($picDivs as $div) {
    $item = [];
    // 提取a标签的href
    $aTag = $xpath->query('.//a', $div)->item(0);
    if ($aTag) {
        $item['href'] = $aTag->getAttribute('href');
    }
    // 提取img标签的src
    $imgTag = $xpath->query('.//img', $div)->item(0);
    if ($imgTag) {
        $item['img_src'] = $imgTag->getAttribute('src');
    }
    // 提取span标签的文本
    $spanTag = $xpath->query('.//span', $div)->item(0);
    if ($spanTag) {
        $item['span_text'] = trim($spanTag->nodeValue);
    }
    // 提取括号内的数字
    $textContent = trim($div->nodeValue);
    preg_match('/\((\d+)\)/', $textContent, $matches);
    if (!empty($matches[1])) {
        $item['number'] = $matches[1];
    }
    $results[] = $item;
}

// 输出结果
print_r($results);

运行后会得到结构化的结果数组,每个元素对应一个条目里的4类信息。

方案2:使用修正后的正则表达式

如果一定要用正则,需要编写包含精准捕获组的表达式,匹配每个ti容器内的目标内容:

$html = '<div class="ti"><div class="pic"> <a href="/categories/rr/1.html"><img src="http://www.erty.com/images/440f2d2a.jpg" alt="Ind"> <span>Ind</span></a> (98) </div></div><div class="ti"><div class="pic"> <a href="/categories/ert/1.html"><img src="http://www.erty.com/images/4123d2b.jpg" alt="Wes"> <span>Wes</span></a> (6044) </div></div>';

// 正则表达式,4个捕获组分别对应你需要的4类信息
$pattern = '/<div class="ti"><div class="pic">\s*<a href="([^"]+)">\s*<img src="([^"]+)"[^>]*>\s*<span>([^<]+)<\/span><\/a>\s*\((\d+)\)\s*<\/div><\/div>/';

preg_match_all($pattern, $html, $matches, PREG_SET_ORDER);

$results = [];
foreach ($matches as $match) {
    $results[] = [
        'href' => $match[1],
        'img_src' => $match[2],
        'span_text' => $match[3],
        'number' => $match[4]
    ];
}

// 输出结果
print_r($results);

这个正则通过4个捕获组分别捕获:

  1. ([^"]+):捕获a标签href属性的内容(直到双引号结束)
  2. ([^"]+):捕获img标签src属性的内容
  3. ([^<]+):捕获span标签内的文本(直到<结束)
  4. (\d+):捕获括号内的数字

注意:这个正则依赖HTML格式稳定,如果标签内出现额外的属性、换行或空格,可能需要调整正则中的\s*(匹配任意空白字符)来兼容。

内容的提问来源于stack exchange,提问作者Blaze Mathew

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 06:51:16