You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用C++调用MediaWiki API获取分类页面以提升Python脚本效率?

Hey there! Great call optimizing your MediaWiki bot with C++—I’ve tackled similar performance bottlenecks with Python-based bots before, and switching critical parts to compiled code makes a huge difference. Let’s walk through exactly how to fetch all pages in a category using C++ and the MediaWiki API.

First, let’s recap the core logic: regardless of language, fetching category pages uses the MediaWiki API’s action=query + list=categorymembers endpoint. We’ll need to handle pagination (since large categories return results in batches) and parse JSON responses—all straightforward in C++ with the right libraries.

Step 1: Set Up Dependencies

We’ll use two battle-tested libraries:

  • libcurl: A cross-platform HTTP client for sending API requests. Install via your package manager (e.g., apt-get install libcurl4-openssl-dev on Debian/Ubuntu, or vcpkg on Windows).
  • nlohmann/json: A lightweight, header-only JSON parser—just download the single header file and include it in your project.

Step 2: Full Working Code

Here’s a complete implementation that fetches all pages in a category, handles pagination automatically, and includes basic error handling:

#include <iostream>
#include <string>
#include <vector>
#include <curl/curl.h>
#include <nlohmann/json.hpp>
#include <thread>
#include <chrono>

using json = nlohmann::json;

// Callback to capture HTTP response data from libcurl
size_t WriteCallback(void* contents, size_t size, size_t nmemb, std::string* s) {
    size_t newLength = size * nmemb;
    try {
        s->append(static_cast<char*>(contents), newLength);
    } catch (const std::bad_alloc& e) {
        return 0; // Handle memory allocation failure
    }
    return newLength;
}

// Send a request to the MediaWiki API and return parsed JSON
json sendApiRequest(const std::string& apiUrl, const std::string& params) {
    CURL* curl = curl_easy_init();
    std::string responseBuffer;
    json responseJson;

    if (!curl) {
        std::cerr << "Failed to initialize libcurl" << std::endl;
        return responseJson;
    }

    std::string fullUrl = apiUrl + "?" + params;
    curl_easy_setopt(curl, CURLOPT_URL, fullUrl.c_str());
    curl_easy_setopt(curl, CURLOPT_WRITEFUNCTION, WriteCallback);
    curl_easy_setopt(curl, CURLOPT_WRITEDATA, &responseBuffer);
    curl_easy_setopt(curl, CURLOPT_FOLLOWLOCATION, 1L); // Follow redirects
    curl_easy_setopt(curl, CURLOPT_USERAGENT, "YourBotName/1.0 (your-contact-info)"); // Required by API rules

    CURLcode result = curl_easy_perform(curl);
    if (result != CURLE_OK) {
        std::cerr << "HTTP request failed: " << curl_easy_strerror(result) << std::endl;
    } else {
        try {
            responseJson = json::parse(responseBuffer);
        } catch (const json::parse_error& e) {
            std::cerr << "JSON parse error: " << e.what() << std::endl;
        }
    }

    curl_easy_cleanup(curl);
    return responseJson;
}

// Fetch all pages in a given category, handling pagination
std::vector<std::string> fetchAllCategoryPages(const std::string& apiUrl, const std::string& categoryName) {
    std::vector<std::string> pageTitles;
    std::string continueToken = "";
    const std::string baseParams = "action=query&list=categorymembers&cmtitle=Category:" + categoryName + "&cmlimit=500&format=json";

    while (true) {
        std::string currentParams = baseParams;
        if (!continueToken.empty()) {
            currentParams += "&cmcontinue=" + continueToken;
        }

        json apiResponse = sendApiRequest(apiUrl, currentParams);

        // Check for API errors first
        if (apiResponse.contains("error")) {
            std::cerr << "API Error: " << apiResponse["error"]["info"].get<std::string>() << std::endl;
            break;
        }

        // Extract page titles from the response
        if (apiResponse.contains("query") && apiResponse["query"].contains("categorymembers")) {
            for (const auto& member : apiResponse["query"]["categorymembers"]) {
                if (member.contains("title")) {
                    pageTitles.push_back(member["title"].get<std::string>());
                }
            }
        }

        // Check if there are more pages to fetch
        if (apiResponse.contains("continue") && apiResponse["continue"].contains("cmcontinue")) {
            continueToken = apiResponse["continue"]["cmcontinue"].get<std::string>();
            // Add a small delay to respect API rate limits (adjust as needed)
            std::this_thread::sleep_for(std::chrono::milliseconds(200));
        } else {
            break; // No more pages
        }
    }

    return pageTitles;
}

int main() {
    // Replace these with your Wiki's API URL and target category
    const std::string wikiApiUrl = "https://your-wiki-domain/w/api.php";
    const std::string targetCategory = "Your_Target_Category";

    std::vector<std::string> categoryPages = fetchAllCategoryPages(wikiApiUrl, targetCategory);

    std::cout << "Successfully fetched " << categoryPages.size() << " pages from Category:" << targetCategory << "\n";
    for (const auto& title : categoryPages) {
        std::cout << "- " << title << "\n";
    }

    return 0;
}

Key Details to Keep in Mind

  1. Pagination Handling: The cmcontinue token tells the API to return the next batch of pages. We loop until this token disappears, meaning we’ve retrieved all pages in the category.
  2. Rate Limiting: The 200ms delay ensures we don’t hit MediaWiki’s default rate limit (5 requests per second). Adjust this based on your wiki’s specific rules.
  3. User-Agent: Always set a meaningful user agent—this is required by MediaWiki’s API terms and helps wiki admins identify your bot.
  4. Authentication: If your wiki is private or requires edit access, add login logic first. Use action=login to get a session cookie, then use libcurl’s CURLOPT_COOKIEJAR and CURLOPT_COOKIEFILE to persist cookies across requests.
  5. Compilation: For Linux, compile with:
    g++ -std=c++17 your-bot.cpp -o your-bot -lcurl
    
    Use C++17 or newer to support nlohmann/json’s full functionality.

This implementation will be significantly faster than Python for large categories, as compiled C++ avoids interpreter overhead and has faster JSON parsing/HTTP handling. You can integrate this with your existing Python code via pybind11 if needed, or use it as a standalone component for the most performance-critical parts of your bot.

内容的提问来源于stack exchange,提问作者Markyroson

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 06:29:54