如何用PHP从HTML提取指定属性并更新数据库?技术实现求助
需求与实现方案
需求概述
需要从files/world/目录下的HTML文件中,提取所有<area>标签的coords属性、href参数中的id值(需解码URL编码)以及alt属性;将提取到的数据整理后,更新数据库hello_world表的content列。
相关代码片段
HTML示例代码
<map id="qmap" name="qmap"> <area shape="rect" coords="430,255,665,304" href="abc.php?qmap=Good World&id=%25SSP%2550469%25" onclick="if( parent.document.body.rows == '*,0' ) parent.document.body.rows = '25%,75%'" alt="SSP 50469" target="content" /> <area shape="rect" coords="457,97,691,116" href="abc.php?qmap=Good World&id=%25SSP%2541175-02%25" onclick="if( parent.document.body.rows == '*,0' ) parent.document.body.rows = '25%,75%'" alt="SSP 41175 Bk 2" target="content" /> </map>
现有PHP代码片段
if (condition) { $html = file_get_contents('files/world/'.$world.'.html'); // Line A // parse the HTML, extract the data, and update the database with the information found in the HTML file (Line B) }
数据库更新参考语句
$db->query("UPDATE `hello_world` SET `content` = '' WHERE `id` = '$file_id'");
Line B 完善实现
推荐使用PHP内置的DOMDocument解析HTML(避免正则解析的兼容性问题),同时用预处理语句防止SQL注入,具体代码如下:
if (condition) { $html = file_get_contents('files/world/'.$world.'.html'); // Line A // Line B 开始 libxml_use_internal_errors(true); // 禁用HTML解析错误提示 $dom = new DOMDocument(); $dom->loadHTML($html); libxml_clear_errors(); $areas = $dom->getElementsByTagName('area'); $extractedData = []; foreach ($areas as $area) { // 提取coords属性 $coords = $area->getAttribute('coords'); // 提取并解析href中的id参数 $href = $area->getAttribute('href'); parse_str(parse_url($href, PHP_URL_QUERY), $queryParams); $id = isset($queryParams['id']) ? urldecode($queryParams['id']) : ''; // 提取alt属性 $alt = $area->getAttribute('alt'); $extractedData[] = [ 'coords' => $coords, 'id' => $id, 'alt' => $alt ]; } // 将提取的数据整理成适合存入content列的格式(示例转为JSON) $content = json_encode($extractedData, JSON_UNESCAPED_UNICODE); // 使用预处理语句更新数据库,防止SQL注入 $stmt = $db->prepare("UPDATE `hello_world` SET `content` = ? WHERE `id` = ?"); $stmt->bind_param("ss", $content, $file_id); $stmt->execute(); $stmt->close(); // Line B 结束 }
代码说明
- HTML解析:用
DOMDocument加载HTML,通过getElementsByTagName获取所有<area>标签,比正则解析更稳定可靠。 - 数据提取:
- 直接通过
getAttribute获取coords和alt属性; - 对
href先解析URL参数,再对id值做URL解码(处理示例中的%25等编码);
- 直接通过
- 数据库更新:用预处理语句绑定参数,避免SQL注入风险;提取的数据转为JSON格式存入
content列(可根据实际需求调整存储格式)。
内容的提问来源于stack exchange,提问作者flash
相关产品推荐
相关产品推荐

