You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用preg_match或其他方法将HTML数据提取到PHP数组

提取HTML场所信息为PHP数组的解决方案

直接用正则处理HTML容易因格式细节失效,更可靠的方式是使用PHP的DOM解析器来处理,以下是完整实现代码:

$html = '<p><b>Ado’s Kitchen &amp; Bar&nbsp; </b>1143 13th St., 720-465-9063; <a href="http://www.span-ishatthehill.com">span-ishatthehill.com.</a> Laid back restaurant with global menu. Open for breakfast and lunch daily and dinner Mon.-Sat.</p>

<p><strong>Blackbelly Market</strong> 1606 Conestoga St. #3, 303-247-1000; <a href="http://www.blackbelly.com">blackbelly.com</a>. Locavore dining, butchery and bar. Open daily for happy hour and dinner; see website for market hours.</p>';

$dom = new DOMDocument();
// 忽略旧网站可能存在的HTML格式错误
libxml_use_internal_errors(true);
$dom->loadHTML($html);
libxml_clear_errors();

$places = [];

foreach ($dom->getElementsByTagName('p') as $p) {
    $place = [
        'name' => '',
        'address' => '',
        'phone' => '',
        'website' => '',
        'description' => ''
    ];

    // 提取名称(匹配<b>或<strong>标签)
    foreach ($p->childNodes as $node) {
        if ($node->nodeName === 'b' || $node->nodeName === 'strong') {
            // 清理多余空格和转义字符
            $place['name'] = trim(preg_replace('/\s+/', ' ', htmlspecialchars_decode($node->nodeValue)));
            break;
        }
    }

    // 提取网站URL
    $aNode = $p->getElementsByTagName('a')->item(0);
    if ($aNode) {
        $place['website'] = $aNode->getAttribute('href');
    }

    // 处理剩余文本,拆分地址、电话和描述
    $fullText = trim(preg_replace('/\s+/', ' ', htmlspecialchars_decode($p->textContent)));
    // 移除名称部分
    $textWithoutName = preg_replace('/^' . preg_quote($place['name'], '/') . '/', '', $fullText);
    $textWithoutName = trim($textWithoutName);

    // 拆分地址(到第一个逗号位置)
    $commaPos = strpos($textWithoutName, ',');
    if ($commaPos !== false) {
        $place['address'] = trim(substr($textWithoutName, 0, $commaPos));
        $remaining = trim(substr($textWithoutName, $commaPos + 1));

        // 拆分电话(到第一个分号位置)
        $semicolonPos = strpos($remaining, ';');
        if ($semicolonPos !== false) {
            $place['phone'] = trim(substr($remaining, 0, $semicolonPos));
            $remainingAfterPhone = trim(substr($remaining, $semicolonPos + 1));

            // 移除网站显示文本,剩余内容为描述
            if ($aNode) {
                $websiteText = trim(htmlspecialchars_decode($aNode->nodeValue));
                $remainingAfterPhone = preg_replace('/^' . preg_quote($websiteText, '/') . '\.?/', '', $remainingAfterPhone);
            }
            $place['description'] = trim($remainingAfterPhone);
        }
    }

    $places[] = $place;
}

// 输出结果
print_r($places);

代码说明

  1. DOM解析器:相比正则,能更好处理HTML标签的变化(比如同时支持<b>和<strong>),兼容旧网站不规范的HTML格式。
  2. 名称提取:遍历<p>内的子节点,匹配加粗标签,清理文本中的多余空格和转义字符(比如&nbsp;)。
  3. 网站提取:直接获取<a>标签的href属性,确保拿到完整网址。
  4. 地址/电话/描述拆分:通过标点符号(逗号、分号)作为分隔符,拆分出对应内容,最后移除网站显示文本得到描述。

如果旧网站的HTML格式有轻微变动(比如标点差异),只需微调拆分逻辑即可,整体稳定性远高于正则方案。

内容的提问来源于stack exchange,提问作者Yogesh Saroya

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 09:02:51