You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用PHP从HTML提取指定属性并更新数据库?技术实现求助

需求与实现方案

需求概述

需要从files/world/目录下的HTML文件中,提取所有<area>标签的coords属性、href参数中的id值(需解码URL编码)以及alt属性;将提取到的数据整理后,更新数据库hello_world表的content列。

相关代码片段

HTML示例代码

<map id="qmap" name="qmap">
    <area shape="rect" coords="430,255,665,304" href="abc.php?qmap=Good World&id=%25SSP%2550469%25" onclick="if( parent.document.body.rows == '*,0' ) parent.document.body.rows = '25%,75%'" alt="SSP 50469" target="content" />
    <area shape="rect" coords="457,97,691,116" href="abc.php?qmap=Good World&id=%25SSP%2541175-02%25" onclick="if( parent.document.body.rows == '*,0' ) parent.document.body.rows = '25%,75%'" alt="SSP 41175 Bk 2" target="content" />
</map>

现有PHP代码片段

if (condition) {
    $html = file_get_contents('files/world/'.$world.'.html'); // Line A
    // parse the HTML, extract the data, and update the database with the information found in the HTML file (Line B)
}

数据库更新参考语句

$db->query("UPDATE `hello_world` SET `content` = '' WHERE `id` = '$file_id'");

Line B 完善实现

推荐使用PHP内置的DOMDocument解析HTML(避免正则解析的兼容性问题),同时用预处理语句防止SQL注入,具体代码如下:

if (condition) {
    $html = file_get_contents('files/world/'.$world.'.html'); // Line A
    
    // Line B 开始
    libxml_use_internal_errors(true); // 禁用HTML解析错误提示
    $dom = new DOMDocument();
    $dom->loadHTML($html);
    libxml_clear_errors();

    $areas = $dom->getElementsByTagName('area');
    $extractedData = [];

    foreach ($areas as $area) {
        // 提取coords属性
        $coords = $area->getAttribute('coords');
        // 提取并解析href中的id参数
        $href = $area->getAttribute('href');
        parse_str(parse_url($href, PHP_URL_QUERY), $queryParams);
        $id = isset($queryParams['id']) ? urldecode($queryParams['id']) : '';
        // 提取alt属性
        $alt = $area->getAttribute('alt');

        $extractedData[] = [
            'coords' => $coords,
            'id' => $id,
            'alt' => $alt
        ];
    }

    // 将提取的数据整理成适合存入content列的格式(示例转为JSON)
    $content = json_encode($extractedData, JSON_UNESCAPED_UNICODE);

    // 使用预处理语句更新数据库,防止SQL注入
    $stmt = $db->prepare("UPDATE `hello_world` SET `content` = ? WHERE `id` = ?");
    $stmt->bind_param("ss", $content, $file_id);
    $stmt->execute();
    $stmt->close();
    // Line B 结束
}

代码说明

  1. HTML解析:用DOMDocument加载HTML,通过getElementsByTagName获取所有<area>标签,比正则解析更稳定可靠。
  2. 数据提取:
    • 直接通过getAttribute获取coords和alt属性;
    • 对href先解析URL参数,再对id值做URL解码(处理示例中的%25等编码);
  3. 数据库更新:用预处理语句绑定参数,避免SQL注入风险;提取的数据转为JSON格式存入content列(可根据实际需求调整存储格式)。

内容的提问来源于stack exchange,提问作者flash

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 14:02:52