如何将cl-html5-parser解析的列表格式HTML转回HTML?
HTML解析与序列化的Lisp工具适配问题
格式差异原因
你遇到的报错核心是cl-html5-parser的:xmls模式输出与XMLS的预期结构不兼容:
- XMLS的
write-xml要求输入是XMLS:NODE结构体、数字、符号或字符串; - 但cl-html5-parser输出的是纯嵌套列表(比如
("html" (("lang" "en")) ...)),且对空属性/节点用#标记,这种格式无法被XMLS的类型匹配逻辑识别,导致ETYPECASE分支匹配失败。
你的测试代码:
(xmls:write-xml ;take the parsed html and convert back to html (html5-parser:parse-html5 ;parse the html into a list (dex:get "https://example.com/") :dom :xmls) ;get the html (open "text2.txt") ; stream, required by xmls )
报错核心片段:
debugger invoked on a SB-KERNEL:CASE-FAILURE @B8012CA9EA in thread #<THREAD tid=358582 "main thread" RUNNING {1200BE8143}>: ("html" (("lang" "en")) ("head" NIL ("title" NIL "Example Domain") ("meta" (# #)) ("style" NIL "body{background:#eee;width:60vw;margin:15vh auto;font-family:system-ui,sans-serif}h1{font-size:1.5em}div{opacity:0.8}a:link,a:visited{color:#348}")) ("body" NIL ("div" NIL ("h1" NIL "Example Domain") ("p" NIL "This domain is for use in documentation examples without needing permission. Avoid use in operations.") ("p" NIL #)) " ")) fell through ETYPECASE expression. Wanted one of (XMLS:NODE NUMBER SYMBOL STRING).
Lisp中HTML列表结构的标准情况
Lisp生态里没有统一的HTML列表结构标准,不同库会自定义格式:
- XMLS使用带类型的节点结构体(
XMLS:NODE),包含名称、属性、子节点等字段; - cl-html5-parser的
:xmls模式只是模仿XMLS的列表结构,但细节处理(比如空节点、属性的表示)不一样; - 其他库(如Plump)则使用自己的对象模型而非纯列表。
推荐的完整流程工具
1. Plump + Plump-Sequencer
这是Lisp中处理HTML最成熟的组合之一:
- 解析:用
plump:parse将HTML字符串解析为可操作的节点对象; - 修改:直接操作节点对象(比如修改文本、添加/删除元素);
- 序列化:用
plump-sequencer:serialize将修改后的节点树输出为HTML字符串。
示例代码:
(let* ((html (dex:get "https://example.com/")) (document (plump:parse html))) ;; 操作节点:修改h1标签文本 (setf (plump:text (first (plump:get-elements-by-tag document "h1"))) "Modified Example Domain") ;; 序列化为HTML并写入文件 (with-open-file (stream "modified.html" :direction :output :if-exists :supersede) (plump-sequencer:serialize document stream)))
2. cl-libxml2
如果你需要更贴近标准XML/HTML的处理,可以使用libxml2的Lisp绑定:
- 支持完整的HTML5解析与序列化,结构规范;
- 适合复杂的DOM操作,比如XPath查询、节点批量修改等。
3. 手动适配cl-html5-parser与XMLS(不推荐)
如果坚持使用现有工具,需要手动转换格式:
- 把cl-html5-parser输出中的
#替换为空列表或空字符串; - 将纯列表结构转换为XMLS的
NODE实例(用xmls:make-node函数)。
但这种方式需要处理大量细节,维护成本高,不如直接使用配套工具。
内容的提问来源于stack exchange,提问作者Oliver Cox
相关产品推荐
相关产品推荐

