You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用C++解析Apache目录列表HTML表格中的JSON文件链接?

问题描述

需要用C++解析Apache目录列表的HTML代码,提取表格行中所有JSON文件的相对URL,最终得到包含3个JSON文件路径的std::vector<std::string>。目标HTML代码如下:

<!DOCTYPE html>
<html>
    <head>
        <meta http-equiv="Content-type" content="text/html; charset=UTF-8"/>
        <meta name="viewport" content="width=device-width, initial-scale=1.0"/>
        <link rel="stylesheet" href="/_autoindex/assets/css/autoindex.css"/>
        <script src="/_autoindex/assets/js/tablesort.js"></script>
        <script src="/_autoindex/assets/js/tablesort.number.js"></script>
        <title>Index of /mydirectory/subdirectory/</title>
    </head>
    <body>
        <div class="content">
            <h1>Index of /mydirectory/subdirectory/</h1>
            <div id="table-list">
                <table id="table-content">
                    <thead class="t-header">
                        <tr>
                            <th class="colname" aria-sort="ascending">
                                <a class="name" href="?ND" onclick="return false"">Name</a></th><th class=" colname " data-sort-method=" number "><a href=" ?MA "  onclick=" return false"">Last Modified</a>
                            </th>
                            <th class="colname" data-sort-method="number"><a href="?SA"onclick="return false"">Size</a></th></tr></thead>
<tr data-sort-method="none "><td><a href="/mydirectory/"><img class="icon " src="/_autoindex/assets/icons/corner-left-up.svg " alt="Up ">Parent Directory</a></td><td></td><td></td></tr>
<tr><td data-sort="first.json "><a href="/mydirectory/subdirectory/first.json "><img class="icon " src="/_autoindex/assets/icons/file.svg " alt="File ">first.json</a></td><td data-sort="1704288747 ">2024-01-03 13:32</td><td data-sort="4096 ">      4k</td></tr>
<tr><td data-sort="second.json "><a href="/mydirectory/subdirectory/second.json "><img class="icon " src="/_autoindex/assets/icons/file.svg " alt="File ">second.json</a></td><td data-sort="1704290309 ">2024-01-03 13:58</td><td data-sort="4096 ">      4k</td></tr>
<tr><td data-sort="third.json "><a href="/mydirectory/subdirectory/third.json "><img class="icon " src="/_autoindex/assets/icons/file.svg " alt="File ">third.json</a></td><td data-sort="1704290300 ">2024-01-03 13:58</td><td data-sort="4096 ">      4k</td></tr>
</table></div>
<address>Proudly Served by LiteSpeed Web Server at example.com Port 443</address></div><script>new Tablesort(document.getElementById("table-content "));</script></body></html>

尝试使用Apache Xerces-C库实现,但该库缺乏完整XPath支持;原本计划用的Xalan-C库无法通过vcpkg获取。现寻求:

  1. 如何用Xerces-C实现类似JSoup的HTML解析功能,完成需求?
  2. 推荐带有vcpkg端口的其他HTML解析库。

已有部分实现代码:

std::vector<std::string> parse_all_links(const std::string &website_content)
{
    std::vector<std::string> collected_links;
    
    try
    {
        XMLPlatformUtils::Initialize();
    }
    catch (const XMLException& exception)
    {
        auto error_message = XMLString::transcode(exception.getMessage());
        logger->error("Failed to initialize XML platform utils: " + std::string(error_message));
        XMLString::release(&error_message);

        return collected_links;
    }
    
    {   
        XercesDOMParser parser;
        parser.setValidationScheme(XercesDOMParser::Val_Never);

        const MemBufInputSource input_source(reinterpret_cast<const XMLByte*>(website_content.data()),
            website_content.size(), "dummy");
        parser.parse(input_source);

        // ...  
    }
    
    XMLPlatformUtils::Terminate();

    return collected_links;
}
解决方案

一、使用Xerces-C实现解析

由于Xerces-C是XML解析库,处理HTML需要手动遍历DOM树模拟选择逻辑,以下是补充完整的代码:

#include <xercesc/dom/DOM.hpp>
#include <xercesc/util/XMLString.hpp>
#include <xercesc/parsers/XercesDOMParser.hpp>
#include <xercesc/sax/HandlerBase.hpp>
#include <xercesc/util/PlatformUtils.hpp>
#include <vector>
#include <string>
#include <algorithm>

using namespace xercesc;

std::vector<std::string> parse_all_links(const std::string &website_content)
{
    std::vector<std::string> collected_links;
    
    try
    {
        XMLPlatformUtils::Initialize();
    }
    catch (const XMLException& exception)
    {
        auto error_message = XMLString::transcode(exception.getMessage());
        logger->error("Failed to initialize XML platform utils: " + std::string(error_message));
        XMLString::release(&error_message);
        return collected_links;
    }
    
    {   
        XercesDOMParser parser;
        parser.setValidationScheme(XercesDOMParser::Val_Never);
        parser.setDoNamespaces(false);
        parser.setLoadExternalDTD(false);

        const MemBufInputSource input_source(reinterpret_cast<const XMLByte*>(website_content.data()),
            website_content.size(), "dummy");
        
        try {
            parser.parse(input_source);
        } catch (const XMLException& e) {
            auto err = XMLString::transcode(e.getMessage());
            logger->error("Parse error: " + std::string(err));
            XMLString::release(&err);
            XMLPlatformUtils::Terminate();
            return collected_links;
        } catch (const DOMException& e) {
            auto err = XMLString::transcode(e.getMessage());
            logger->error("DOM error: " + std::string(err));
            XMLString::release(&err);
            XMLPlatformUtils::Terminate();
            return collected_links;
        }

        DOMDocument* doc = parser.getDocument();
        if (!doc) {
            XMLPlatformUtils::Terminate();
            return collected_links;
        }

        DOMElement* root = doc->getDocumentElement();
        DOMNodeList* tr_list = root->getElementsByTagName(XMLString::transcode("tr"));
        if (!tr_list) {
            XMLPlatformUtils::Terminate();
            return collected_links;
        }

        const XMLSize_t tr_count = tr_list->getLength();
        for (XMLSize_t i = 0; i < tr_count; ++i) {
            DOMNode* tr_node = tr_list->item(i);
            if (tr_node->getNodeType() != DOMNode::ELEMENT_NODE) {
                continue;
            }
            DOMElement* tr_elem = static_cast<DOMElement*>(tr_node);

            DOMNodeList* td_list = tr_elem->getElementsByTagName(XMLString::transcode("td"));
            if (!td_list || td_list->getLength() == 0) {
                continue;
            }

            DOMElement* first_td = static_cast<DOMElement*>(td_list->item(0));
            DOMNodeList* a_list = first_td->getElementsByTagName(XMLString::transcode("a"));
            if (!a_list || a_list->getLength() == 0) {
                continue;
            }

            DOMElement* a_elem = static_cast<DOMElement*>(a_list->item(0));
            const XMLCh* href_attr = a_elem->getAttribute(XMLString::transcode("href"));
            if (!href_attr) {
                continue;
            }

            char* href_cstr = XMLString::transcode(href_attr);
            std::string href(href_cstr);
            XMLString::release(&href_cstr);

            if (href.find(".json") != std::string::npos && href != "/mydirectory/") {
                const std::string base_path = "/mydirectory/subdirectory/";
                size_t pos = href.find(base_path);
                if (pos != std::string::npos) {
                    std::string relative_url = href.substr(pos + base_path.length());
                    relative_url.erase(std::remove_if(relative_url.begin(), relative_url.end(), isspace), relative_url.end());
                    collected_links.push_back(relative_url);
                }
            }
        }

        tr_list->release();
        parser.resetDocumentPool();
    }
    
    XMLPlatformUtils::Terminate();
    return collected_links;
}

代码说明

  1. 解析器配置:关闭验证、命名空间处理和外部DTD加载,适配HTML的松散格式。
  2. DOM遍历:逐层查找<tr>→<td>→<a>节点,提取href属性值。
  3. 链接过滤:筛选JSON文件链接,排除父目录链接,提取并清理相对URL。
  4. 资源管理:手动释放Xerces-C分配的字符串和节点列表,避免内存泄漏。

二、推荐带vcpkg端口的HTML解析库

以下是几个更适合HTML解析、可通过vcpkg安装的库:

1. Gumbo Parser

  • 特点:纯C编写的HTML5解析器,输出标准DOM树,支持容错解析不规范HTML,API简洁。
  • vcpkg安装命令:vcpkg install gumbo
  • 优势:轻量高效,专门为HTML设计,比XML解析库更适配网页内容。

2. libxml2

  • 特点:成熟的XML/HTML解析库,支持XPath查询,启用HTML模式后可处理网页内容。
  • vcpkg安装命令:vcpkg install libxml2
  • 优势:支持XPath,能简化复杂元素选择逻辑,适合需要精准查询的场景。

3. pugixml

  • 特点:轻量级XML/HTML解析库,API简单易用,支持容错解析,性能优异。
  • vcpkg安装命令:vcpkg install pugixml
  • 优势:体积小,集成方便,适合嵌入式或资源受限场景,能处理大多数HTML内容。

4. cpprestsdk (Casablanca)

  • 特点:微软开源的C++ REST库,内置HTML解析模块,同时支持HTTP请求、JSON处理等功能。
  • vcpkg安装命令:vcpkg install cpprestsdk
  • 优势:如果项目需要HTTP请求+HTML解析,可一站式解决,API现代化。

内容的提问来源于stack exchange,提问作者BullyWiiPlaza

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 22:04:57