如何使用preg_match或其他方法将HTML数据提取到PHP数组
提取HTML场所信息为PHP数组的解决方案
直接用正则处理HTML容易因格式细节失效,更可靠的方式是使用PHP的DOM解析器来处理,以下是完整实现代码:
$html = '<p><b>Ado’s Kitchen & Bar </b>1143 13th St., 720-465-9063; <a href="http://www.span-ishatthehill.com">span-ishatthehill.com.</a> Laid back restaurant with global menu. Open for breakfast and lunch daily and dinner Mon.-Sat.</p> <p><strong>Blackbelly Market</strong> 1606 Conestoga St. #3, 303-247-1000; <a href="http://www.blackbelly.com">blackbelly.com</a>. Locavore dining, butchery and bar. Open daily for happy hour and dinner; see website for market hours.</p>'; $dom = new DOMDocument(); // 忽略旧网站可能存在的HTML格式错误 libxml_use_internal_errors(true); $dom->loadHTML($html); libxml_clear_errors(); $places = []; foreach ($dom->getElementsByTagName('p') as $p) { $place = [ 'name' => '', 'address' => '', 'phone' => '', 'website' => '', 'description' => '' ]; // 提取名称(匹配<b>或<strong>标签) foreach ($p->childNodes as $node) { if ($node->nodeName === 'b' || $node->nodeName === 'strong') { // 清理多余空格和转义字符 $place['name'] = trim(preg_replace('/\s+/', ' ', htmlspecialchars_decode($node->nodeValue))); break; } } // 提取网站URL $aNode = $p->getElementsByTagName('a')->item(0); if ($aNode) { $place['website'] = $aNode->getAttribute('href'); } // 处理剩余文本,拆分地址、电话和描述 $fullText = trim(preg_replace('/\s+/', ' ', htmlspecialchars_decode($p->textContent))); // 移除名称部分 $textWithoutName = preg_replace('/^' . preg_quote($place['name'], '/') . '/', '', $fullText); $textWithoutName = trim($textWithoutName); // 拆分地址(到第一个逗号位置) $commaPos = strpos($textWithoutName, ','); if ($commaPos !== false) { $place['address'] = trim(substr($textWithoutName, 0, $commaPos)); $remaining = trim(substr($textWithoutName, $commaPos + 1)); // 拆分电话(到第一个分号位置) $semicolonPos = strpos($remaining, ';'); if ($semicolonPos !== false) { $place['phone'] = trim(substr($remaining, 0, $semicolonPos)); $remainingAfterPhone = trim(substr($remaining, $semicolonPos + 1)); // 移除网站显示文本,剩余内容为描述 if ($aNode) { $websiteText = trim(htmlspecialchars_decode($aNode->nodeValue)); $remainingAfterPhone = preg_replace('/^' . preg_quote($websiteText, '/') . '\.?/', '', $remainingAfterPhone); } $place['description'] = trim($remainingAfterPhone); } } $places[] = $place; } // 输出结果 print_r($places);
代码说明
- DOM解析器:相比正则,能更好处理HTML标签的变化(比如同时支持
<b>和<strong>),兼容旧网站不规范的HTML格式。 - 名称提取:遍历
<p>内的子节点,匹配加粗标签,清理文本中的多余空格和转义字符(比如 )。 - 网站提取:直接获取
<a>标签的href属性,确保拿到完整网址。 - 地址/电话/描述拆分:通过标点符号(逗号、分号)作为分隔符,拆分出对应内容,最后移除网站显示文本得到描述。
如果旧网站的HTML格式有轻微变动(比如标点差异),只需微调拆分逻辑即可,整体稳定性远高于正则方案。
内容的提问来源于stack exchange,提问作者Yogesh Saroya
相关产品推荐
相关产品推荐

