如何用C++调用MediaWiki API获取分类页面以提升Python脚本效率?
Hey there! Great call optimizing your MediaWiki bot with C++—I’ve tackled similar performance bottlenecks with Python-based bots before, and switching critical parts to compiled code makes a huge difference. Let’s walk through exactly how to fetch all pages in a category using C++ and the MediaWiki API.
First, let’s recap the core logic: regardless of language, fetching category pages uses the MediaWiki API’s action=query + list=categorymembers endpoint. We’ll need to handle pagination (since large categories return results in batches) and parse JSON responses—all straightforward in C++ with the right libraries.
Step 1: Set Up Dependencies
We’ll use two battle-tested libraries:
- libcurl: A cross-platform HTTP client for sending API requests. Install via your package manager (e.g.,
apt-get install libcurl4-openssl-devon Debian/Ubuntu, or vcpkg on Windows). - nlohmann/json: A lightweight, header-only JSON parser—just download the single header file and include it in your project.
Step 2: Full Working Code
Here’s a complete implementation that fetches all pages in a category, handles pagination automatically, and includes basic error handling:
#include <iostream> #include <string> #include <vector> #include <curl/curl.h> #include <nlohmann/json.hpp> #include <thread> #include <chrono> using json = nlohmann::json; // Callback to capture HTTP response data from libcurl size_t WriteCallback(void* contents, size_t size, size_t nmemb, std::string* s) { size_t newLength = size * nmemb; try { s->append(static_cast<char*>(contents), newLength); } catch (const std::bad_alloc& e) { return 0; // Handle memory allocation failure } return newLength; } // Send a request to the MediaWiki API and return parsed JSON json sendApiRequest(const std::string& apiUrl, const std::string& params) { CURL* curl = curl_easy_init(); std::string responseBuffer; json responseJson; if (!curl) { std::cerr << "Failed to initialize libcurl" << std::endl; return responseJson; } std::string fullUrl = apiUrl + "?" + params; curl_easy_setopt(curl, CURLOPT_URL, fullUrl.c_str()); curl_easy_setopt(curl, CURLOPT_WRITEFUNCTION, WriteCallback); curl_easy_setopt(curl, CURLOPT_WRITEDATA, &responseBuffer); curl_easy_setopt(curl, CURLOPT_FOLLOWLOCATION, 1L); // Follow redirects curl_easy_setopt(curl, CURLOPT_USERAGENT, "YourBotName/1.0 (your-contact-info)"); // Required by API rules CURLcode result = curl_easy_perform(curl); if (result != CURLE_OK) { std::cerr << "HTTP request failed: " << curl_easy_strerror(result) << std::endl; } else { try { responseJson = json::parse(responseBuffer); } catch (const json::parse_error& e) { std::cerr << "JSON parse error: " << e.what() << std::endl; } } curl_easy_cleanup(curl); return responseJson; } // Fetch all pages in a given category, handling pagination std::vector<std::string> fetchAllCategoryPages(const std::string& apiUrl, const std::string& categoryName) { std::vector<std::string> pageTitles; std::string continueToken = ""; const std::string baseParams = "action=query&list=categorymembers&cmtitle=Category:" + categoryName + "&cmlimit=500&format=json"; while (true) { std::string currentParams = baseParams; if (!continueToken.empty()) { currentParams += "&cmcontinue=" + continueToken; } json apiResponse = sendApiRequest(apiUrl, currentParams); // Check for API errors first if (apiResponse.contains("error")) { std::cerr << "API Error: " << apiResponse["error"]["info"].get<std::string>() << std::endl; break; } // Extract page titles from the response if (apiResponse.contains("query") && apiResponse["query"].contains("categorymembers")) { for (const auto& member : apiResponse["query"]["categorymembers"]) { if (member.contains("title")) { pageTitles.push_back(member["title"].get<std::string>()); } } } // Check if there are more pages to fetch if (apiResponse.contains("continue") && apiResponse["continue"].contains("cmcontinue")) { continueToken = apiResponse["continue"]["cmcontinue"].get<std::string>(); // Add a small delay to respect API rate limits (adjust as needed) std::this_thread::sleep_for(std::chrono::milliseconds(200)); } else { break; // No more pages } } return pageTitles; } int main() { // Replace these with your Wiki's API URL and target category const std::string wikiApiUrl = "https://your-wiki-domain/w/api.php"; const std::string targetCategory = "Your_Target_Category"; std::vector<std::string> categoryPages = fetchAllCategoryPages(wikiApiUrl, targetCategory); std::cout << "Successfully fetched " << categoryPages.size() << " pages from Category:" << targetCategory << "\n"; for (const auto& title : categoryPages) { std::cout << "- " << title << "\n"; } return 0; }
Key Details to Keep in Mind
- Pagination Handling: The
cmcontinuetoken tells the API to return the next batch of pages. We loop until this token disappears, meaning we’ve retrieved all pages in the category. - Rate Limiting: The 200ms delay ensures we don’t hit MediaWiki’s default rate limit (5 requests per second). Adjust this based on your wiki’s specific rules.
- User-Agent: Always set a meaningful user agent—this is required by MediaWiki’s API terms and helps wiki admins identify your bot.
- Authentication: If your wiki is private or requires edit access, add login logic first. Use
action=loginto get a session cookie, then use libcurl’sCURLOPT_COOKIEJARandCURLOPT_COOKIEFILEto persist cookies across requests. - Compilation: For Linux, compile with:
Use C++17 or newer to support nlohmann/json’s full functionality.g++ -std=c++17 your-bot.cpp -o your-bot -lcurl
This implementation will be significantly faster than Python for large categories, as compiled C++ avoids interpreter overhead and has faster JSON parsing/HTTP handling. You can integrate this with your existing Python code via pybind11 if needed, or use it as a standalone component for the most performance-critical parts of your bot.
内容的提问来源于stack exchange,提问作者Markyroson

