如何用C++解析Apache目录列表HTML表格中的JSON文件链接?
问题描述
需要用C++解析Apache目录列表的HTML代码,提取表格行中所有JSON文件的相对URL,最终得到包含3个JSON文件路径的std::vector<std::string>。目标HTML代码如下:
<!DOCTYPE html> <html> <head> <meta http-equiv="Content-type" content="text/html; charset=UTF-8"/> <meta name="viewport" content="width=device-width, initial-scale=1.0"/> <link rel="stylesheet" href="/_autoindex/assets/css/autoindex.css"/> <script src="/_autoindex/assets/js/tablesort.js"></script> <script src="/_autoindex/assets/js/tablesort.number.js"></script> <title>Index of /mydirectory/subdirectory/</title> </head> <body> <div class="content"> <h1>Index of /mydirectory/subdirectory/</h1> <div id="table-list"> <table id="table-content"> <thead class="t-header"> <tr> <th class="colname" aria-sort="ascending"> <a class="name" href="?ND" onclick="return false"">Name</a></th><th class=" colname " data-sort-method=" number "><a href=" ?MA " onclick=" return false"">Last Modified</a> </th> <th class="colname" data-sort-method="number"><a href="?SA"onclick="return false"">Size</a></th></tr></thead> <tr data-sort-method="none "><td><a href="/mydirectory/"><img class="icon " src="/_autoindex/assets/icons/corner-left-up.svg " alt="Up ">Parent Directory</a></td><td></td><td></td></tr> <tr><td data-sort="first.json "><a href="/mydirectory/subdirectory/first.json "><img class="icon " src="/_autoindex/assets/icons/file.svg " alt="File ">first.json</a></td><td data-sort="1704288747 ">2024-01-03 13:32</td><td data-sort="4096 "> 4k</td></tr> <tr><td data-sort="second.json "><a href="/mydirectory/subdirectory/second.json "><img class="icon " src="/_autoindex/assets/icons/file.svg " alt="File ">second.json</a></td><td data-sort="1704290309 ">2024-01-03 13:58</td><td data-sort="4096 "> 4k</td></tr> <tr><td data-sort="third.json "><a href="/mydirectory/subdirectory/third.json "><img class="icon " src="/_autoindex/assets/icons/file.svg " alt="File ">third.json</a></td><td data-sort="1704290300 ">2024-01-03 13:58</td><td data-sort="4096 "> 4k</td></tr> </table></div> <address>Proudly Served by LiteSpeed Web Server at example.com Port 443</address></div><script>new Tablesort(document.getElementById("table-content "));</script></body></html>
尝试使用Apache Xerces-C库实现,但该库缺乏完整XPath支持;原本计划用的Xalan-C库无法通过vcpkg获取。现寻求:
- 如何用Xerces-C实现类似JSoup的HTML解析功能,完成需求?
- 推荐带有vcpkg端口的其他HTML解析库。
已有部分实现代码:
std::vector<std::string> parse_all_links(const std::string &website_content) { std::vector<std::string> collected_links; try { XMLPlatformUtils::Initialize(); } catch (const XMLException& exception) { auto error_message = XMLString::transcode(exception.getMessage()); logger->error("Failed to initialize XML platform utils: " + std::string(error_message)); XMLString::release(&error_message); return collected_links; } { XercesDOMParser parser; parser.setValidationScheme(XercesDOMParser::Val_Never); const MemBufInputSource input_source(reinterpret_cast<const XMLByte*>(website_content.data()), website_content.size(), "dummy"); parser.parse(input_source); // ... } XMLPlatformUtils::Terminate(); return collected_links; }
解决方案
一、使用Xerces-C实现解析
由于Xerces-C是XML解析库,处理HTML需要手动遍历DOM树模拟选择逻辑,以下是补充完整的代码:
#include <xercesc/dom/DOM.hpp> #include <xercesc/util/XMLString.hpp> #include <xercesc/parsers/XercesDOMParser.hpp> #include <xercesc/sax/HandlerBase.hpp> #include <xercesc/util/PlatformUtils.hpp> #include <vector> #include <string> #include <algorithm> using namespace xercesc; std::vector<std::string> parse_all_links(const std::string &website_content) { std::vector<std::string> collected_links; try { XMLPlatformUtils::Initialize(); } catch (const XMLException& exception) { auto error_message = XMLString::transcode(exception.getMessage()); logger->error("Failed to initialize XML platform utils: " + std::string(error_message)); XMLString::release(&error_message); return collected_links; } { XercesDOMParser parser; parser.setValidationScheme(XercesDOMParser::Val_Never); parser.setDoNamespaces(false); parser.setLoadExternalDTD(false); const MemBufInputSource input_source(reinterpret_cast<const XMLByte*>(website_content.data()), website_content.size(), "dummy"); try { parser.parse(input_source); } catch (const XMLException& e) { auto err = XMLString::transcode(e.getMessage()); logger->error("Parse error: " + std::string(err)); XMLString::release(&err); XMLPlatformUtils::Terminate(); return collected_links; } catch (const DOMException& e) { auto err = XMLString::transcode(e.getMessage()); logger->error("DOM error: " + std::string(err)); XMLString::release(&err); XMLPlatformUtils::Terminate(); return collected_links; } DOMDocument* doc = parser.getDocument(); if (!doc) { XMLPlatformUtils::Terminate(); return collected_links; } DOMElement* root = doc->getDocumentElement(); DOMNodeList* tr_list = root->getElementsByTagName(XMLString::transcode("tr")); if (!tr_list) { XMLPlatformUtils::Terminate(); return collected_links; } const XMLSize_t tr_count = tr_list->getLength(); for (XMLSize_t i = 0; i < tr_count; ++i) { DOMNode* tr_node = tr_list->item(i); if (tr_node->getNodeType() != DOMNode::ELEMENT_NODE) { continue; } DOMElement* tr_elem = static_cast<DOMElement*>(tr_node); DOMNodeList* td_list = tr_elem->getElementsByTagName(XMLString::transcode("td")); if (!td_list || td_list->getLength() == 0) { continue; } DOMElement* first_td = static_cast<DOMElement*>(td_list->item(0)); DOMNodeList* a_list = first_td->getElementsByTagName(XMLString::transcode("a")); if (!a_list || a_list->getLength() == 0) { continue; } DOMElement* a_elem = static_cast<DOMElement*>(a_list->item(0)); const XMLCh* href_attr = a_elem->getAttribute(XMLString::transcode("href")); if (!href_attr) { continue; } char* href_cstr = XMLString::transcode(href_attr); std::string href(href_cstr); XMLString::release(&href_cstr); if (href.find(".json") != std::string::npos && href != "/mydirectory/") { const std::string base_path = "/mydirectory/subdirectory/"; size_t pos = href.find(base_path); if (pos != std::string::npos) { std::string relative_url = href.substr(pos + base_path.length()); relative_url.erase(std::remove_if(relative_url.begin(), relative_url.end(), isspace), relative_url.end()); collected_links.push_back(relative_url); } } } tr_list->release(); parser.resetDocumentPool(); } XMLPlatformUtils::Terminate(); return collected_links; }
代码说明
- 解析器配置:关闭验证、命名空间处理和外部DTD加载,适配HTML的松散格式。
- DOM遍历:逐层查找
<tr>→<td>→<a>节点,提取href属性值。 - 链接过滤:筛选JSON文件链接,排除父目录链接,提取并清理相对URL。
- 资源管理:手动释放Xerces-C分配的字符串和节点列表,避免内存泄漏。
二、推荐带vcpkg端口的HTML解析库
以下是几个更适合HTML解析、可通过vcpkg安装的库:
1. Gumbo Parser
- 特点:纯C编写的HTML5解析器,输出标准DOM树,支持容错解析不规范HTML,API简洁。
- vcpkg安装命令:
vcpkg install gumbo - 优势:轻量高效,专门为HTML设计,比XML解析库更适配网页内容。
2. libxml2
- 特点:成熟的XML/HTML解析库,支持XPath查询,启用HTML模式后可处理网页内容。
- vcpkg安装命令:
vcpkg install libxml2 - 优势:支持XPath,能简化复杂元素选择逻辑,适合需要精准查询的场景。
3. pugixml
- 特点:轻量级XML/HTML解析库,API简单易用,支持容错解析,性能优异。
- vcpkg安装命令:
vcpkg install pugixml - 优势:体积小,集成方便,适合嵌入式或资源受限场景,能处理大多数HTML内容。
4. cpprestsdk (Casablanca)
- 特点:微软开源的C++ REST库,内置HTML解析模块,同时支持HTTP请求、JSON处理等功能。
- vcpkg安装命令:
vcpkg install cpprestsdk - 优势:如果项目需要HTTP请求+HTML解析,可一站式解决,API现代化。
内容的提问来源于stack exchange,提问作者BullyWiiPlaza
相关产品推荐
相关产品推荐

