在PHP中提取HTML内Jinja语法并转换为可操作标签的方案问询
问题描述
我需要从HTML中提取Jinja语法并将其转换为前端可操作的真实标签。目前用preg_replace_callback实现,但这个方法只对短字符串有效,处理不了包含嵌套条件语句的大段内容。
比如这段包含Jinja条件的HTML:
<span>{% if is_vardiya %}</span><span style="font-size: 10pt"> </span> </p> <p style="margin-top: 0pt; margin-bottom: 0pt; text-align: justify"> <span style="-aw-import: ignore"> </span> </p> <p style=" margin-top: 0pt; margin-bottom: 0pt; text-align: justify; font-size: 12pt; " > <span >Çalışan. </span> </p> <p style=" margin-top: 0pt; margin-bottom: 0pt; text-align: justify; font-size: 12pt; " > <span>{% endif %}</span>
我需要获取{% if is_vardiya %}和{% endif %}之间的内容,同时要支持处理嵌套的if块。我知道正则不是处理HTML和这类模板语法的合适方案,希望能得到更有效的解决方案。
我尝试的PHP代码如下:
$html = preg_replace_callback( '/{%\s*(p if|if)\s*(?:\w+)(?:[\w=!\'. ]*)\s*%}([\w\s+<\/>="-:;])*(?:{%\s*(p if|if)\s*(?:\w+)(?:[\w=!\'. ]*)\s*%}[\w\s+<\/>="-:;]*(?:{%\s*else\s*%}[\w\s+<\/>="-:;]*)*{%\s*(p endif|endif)\s*%}([\w\s+<\/>="-:;])*|(?R))*([\w\s+<\/>="-:;])*(?:{%\s*else\s*%}[\w\s+<\/>="-:;]+)*{%\s*(p endif|endif)\s*%}/', function ($matches) { $matches[0] = preg_replace('/\\r/', '', $matches[0]); $matches[0] = preg_replace('/\\n/', '', $matches[0]); preg_match_all('/{%\s*(p if|if)\s*(?:\w+)(?:[\w=!\'. ]*)\s*%}([\w\s+<\/>="-:;])*{%\s*(p endif|endif)\s*%}/', $matches[0], $subconditions,PREG_OFFSET_CAPTURE); if(count($subconditions[0])>0) { $allSubconditions = $subconditions[0]; if($allSubconditions[0][1] != 0) { // to-do : create recursive function to support infinite if inside if foreach ($allSubconditions as $condition) { $subconditionString = $this->subconditionExpression($condition[0]); $matches[0] = str_replace($condition[0], $subconditionString, $matches[0]); } } } return $this->createMainConditionExpression($matches[0]); }, $html );
解决方案
1. 使用Jinja兼容的模板引擎(推荐)
直接用PHP实现的Jinja兼容引擎(比如Twig,语法和Jinja高度匹配)来处理,它能原生支持嵌套条件解析:
- 将带Jinja标签的HTML作为模板加载,通过引擎的抽象语法树(AST)功能提取条件块信息。
- 解析后可以精准获取每个条件的表达式和内部内容,再转换为前端需要的标签格式。
示例代码思路:
// 初始化Twig环境 $loader = new \Twig\Loader\ArrayLoader(['template' => $html]); $twig = new \Twig\Environment($loader); // 获取模板的AST结构 $tokenStream = $twig->tokenize($html); $ast = $twig->parse($tokenStream); // 遍历AST节点提取条件块 foreach ($ast->getNode('body') as $node) { if ($node instanceof \Twig\Node\IfNode) { // 获取if条件表达式 $conditionExpr = $node->getNode('tests')->getNode(0)->getNode('expr')->getAttribute('value'); // 编译if块内的内容为HTML字符串 $ifContent = $twig->compile($node->getNode('tests')->getNode(0)->getNode('body')); // 这里将条件和内容转换为前端标签,比如生成带data属性的容器 $convertedTag = sprintf('<div data-if="%s">%s</div>', htmlspecialchars($conditionExpr), $ifContent); // 后续可以替换原内容或收集结果 } }
2. 手动编写状态机解析
如果不想依赖第三方库,可以写一个状态机遍历字符串,跟踪Jinja标签的嵌套层级:
- 定义状态:正常文本、识别Jinja标签开始、解析标签内容、识别标签结束。
- 遇到
{% if ... %}时增加嵌套层级并开始记录内容;遇到{% endif %}时减少层级,层级归0时结束当前条件块的记录。 - 这种方法能准确处理嵌套结构,也能兼容复杂HTML。
示例代码:
$html = htmlspecialchars_decode($html); // 先还原HTML实体,避免干扰 $content = $html; $len = strlen($content); $pos = 0; $currentBlock = []; $nestLevel = 0; $inBlock = false; $results = []; while ($pos < $len) { $startTag = strpos($content, '{%', $pos); if ($startTag === false) break; // 记录标签前的内容(如果处于条件块内) if ($startTag > $pos && $inBlock) { $currentBlock['content'] .= substr($content, $pos, $startTag - $pos); } $endTag = strpos($content, '%}', $startTag); if ($endTag === false) break; $tagContent = trim(substr($content, $startTag + 2, $endTag - $startTag - 2)); $pos = $endTag + 2; if (str_starts_with($tagContent, 'if ')) { $nestLevel++; if ($nestLevel == 1) { $inBlock = true; $currentBlock = [ 'condition' => trim(substr($tagContent, 3)), 'content' => '' ]; } elseif ($inBlock) { // 嵌套if,保留标签本身 $currentBlock['content'] .= "{% $tagContent %}"; } } elseif ($tagContent === 'endif') { $nestLevel--; if ($nestLevel == 0) { // 条件块结束,加入结果集 $results[] = $currentBlock; $inBlock = false; } elseif ($inBlock) { // 嵌套endif,保留标签 $currentBlock['content'] .= '{% endif %}'; } } elseif ($inBlock) { // 其他Jinja标签(如else),直接加入内容 $currentBlock['content'] .= "{% $tagContent %}"; } } // 处理结果集,转换为前端标签 foreach ($results as $block) { echo '<div data-condition="' . htmlspecialchars($block['condition']) . '">' . $block['content'] . '</div>'; }
3. 预处理:还原HTML实体
注意到你提供的HTML包含<、"这类实体字符,建议先用htmlspecialchars_decode将其还原为真实的HTML标签,再进行解析,避免实体字符干扰标签匹配逻辑。
内容的提问来源于stack exchange,提问作者Tushar Sharma
相关产品推荐
相关产品推荐

