如何用PHP的preg_match_all提取指定HTML内容中的信息?
如何用PHP的preg_match_all提取指定HTML内容?
需求说明
你需要从以下HTML代码中提取每个条目里的4类信息:
- a标签的href属性值(如
/categories/rr/1.html) - img标签的src属性值(如
http://www.erty.com/images/440f2d2a.jpg) - span标签内的文本(如
Ind) - 括号内的数字(如
98)
用户提供的原始HTML代码:
<div class="ti"><div class="pic"> <a href="/categories/rr/1.html"><img src="http://www.erty.com/images/440f2d2a.jpg" alt="Ind"> <span>Ind</span></a> (98) </div></div><div class="ti"><div class="pic"> <a href="/categories/ert/1.html"><img src="http://www.erty.com/images/4123d2b.jpg" alt="Wes"> <span>Wes</span></a> (6044) </div></div>
原正则的问题
你尝试的正则表达式:
preg_match_all('|[^<div class="ti"><div class="pic">].*?[^</div></div>]+|', $test_html, $out, PREG_PATTERN_ORDER);
这个写法逻辑有误:
[^<div class="ti"><div class="pic">]是否定字符集,它匹配的是不是这些单个字符的内容,而不是排除整个标签字符串,这完全不符合你的预期。- 整个正则没有针对你需要的4个信息设置捕获组,自然无法提取目标内容。
解决方案
方案1:使用DOMDocument(推荐,处理HTML更可靠)
正则处理HTML容易因为标签格式变化(比如空格、换行)失效,PHP的DOM扩展是更专业的HTML解析工具,示例代码如下:
$html = '<div class="ti"><div class="pic"> <a href="/categories/rr/1.html"><img src="http://www.erty.com/images/440f2d2a.jpg" alt="Ind"> <span>Ind</span></a> (98) </div></div><div class="ti"><div class="pic"> <a href="/categories/ert/1.html"><img src="http://www.erty.com/images/4123d2b.jpg" alt="Wes"> <span>Wes</span></a> (6044) </div></div>'; $dom = new DOMDocument(); // 处理HTML可能存在的格式问题 libxml_use_internal_errors(true); $dom->loadHTML($html); libxml_clear_errors(); $xpath = new DOMXPath($dom); // 定位所有class为ti的div下的pic容器 $picDivs = $xpath->query('//div[@class="ti"]/div[@class="pic"]'); $results = []; foreach ($picDivs as $div) { $item = []; // 提取a标签的href $aTag = $xpath->query('.//a', $div)->item(0); if ($aTag) { $item['href'] = $aTag->getAttribute('href'); } // 提取img标签的src $imgTag = $xpath->query('.//img', $div)->item(0); if ($imgTag) { $item['img_src'] = $imgTag->getAttribute('src'); } // 提取span标签的文本 $spanTag = $xpath->query('.//span', $div)->item(0); if ($spanTag) { $item['span_text'] = trim($spanTag->nodeValue); } // 提取括号内的数字 $textContent = trim($div->nodeValue); preg_match('/\((\d+)\)/', $textContent, $matches); if (!empty($matches[1])) { $item['number'] = $matches[1]; } $results[] = $item; } // 输出结果 print_r($results);
运行后会得到结构化的结果数组,每个元素对应一个条目里的4类信息。
方案2:使用修正后的正则表达式
如果一定要用正则,需要编写包含精准捕获组的表达式,匹配每个ti容器内的目标内容:
$html = '<div class="ti"><div class="pic"> <a href="/categories/rr/1.html"><img src="http://www.erty.com/images/440f2d2a.jpg" alt="Ind"> <span>Ind</span></a> (98) </div></div><div class="ti"><div class="pic"> <a href="/categories/ert/1.html"><img src="http://www.erty.com/images/4123d2b.jpg" alt="Wes"> <span>Wes</span></a> (6044) </div></div>'; // 正则表达式,4个捕获组分别对应你需要的4类信息 $pattern = '/<div class="ti"><div class="pic">\s*<a href="([^"]+)">\s*<img src="([^"]+)"[^>]*>\s*<span>([^<]+)<\/span><\/a>\s*\((\d+)\)\s*<\/div><\/div>/'; preg_match_all($pattern, $html, $matches, PREG_SET_ORDER); $results = []; foreach ($matches as $match) { $results[] = [ 'href' => $match[1], 'img_src' => $match[2], 'span_text' => $match[3], 'number' => $match[4] ]; } // 输出结果 print_r($results);
这个正则通过4个捕获组分别捕获:
([^"]+):捕获a标签href属性的内容(直到双引号结束)([^"]+):捕获img标签src属性的内容([^<]+):捕获span标签内的文本(直到<结束)(\d+):捕获括号内的数字
注意:这个正则依赖HTML格式稳定,如果标签内出现额外的属性、换行或空格,可能需要调整正则中的\s*(匹配任意空白字符)来兼容。
内容的提问来源于stack exchange,提问作者Blaze Mathew
相关产品推荐
相关产品推荐

