C++初学者求助:如何用多线程加速跨数据库数据迁移?
Hey there! Let's work through this multi-threading problem together—since you're a C++ beginner, I'll break everything down clearly so you follow along easily. The core goal is to spin up two transfer threads when your loop hits the 10000th element, while letting the main thread keep processing data. First, we need to tackle thread safety to avoid race conditions when accessing vectors from multiple threads.
Step 1: Include Required Headers
Start by adding the standard library headers we need for threads, mutexes, and vectors:
#include <iostream> #include <vector> #include <thread> #include <mutex>
Step 2: Thread-Safe Data Transfer Function
We'll create a helper function that handles copying data from the source vector to the target. We use a mutex to lock access to the source vector during the transfer, so the main thread can't modify it mid-operation:
void transferData(const std::vector<std::vector<int>>& src, std::vector<std::vector<int>>& dest, std::mutex& mtx) { // RAII-style lock: automatically unlocks when we leave this scope std::lock_guard<std::mutex> lock(mtx); // Copy all elements from the source to the target vector dest.insert(dest.end(), src.begin(), src.end()); }
Step 3: Main Program with Thread Logic
Now let's build out the main function. We'll add mutexes to protect our source vectors, start threads at the right time, and ensure we wait for threads to finish before exiting:
int main() { std::vector<std::vector<int>> dataOne; std::vector<std::vector<int>> dataTwo; std::vector<std::vector<int>> dataThree; std::vector<std::vector<int>> dataFour; // Mutexes to prevent race conditions when accessing dataOne and dataTwo std::mutex mtxOne, mtxTwo; // Flag to avoid starting threads multiple times bool threadsStarted = false; const int THRESHOLD = 10000; const int TOTAL_ITERATIONS = 10000000; for (int i = 0; i < TOTAL_ITERATIONS; ++i) { // Lock before modifying dataOne to keep access thread-safe { std::lock_guard<std::mutex> lock(mtxOne); dataOne.push_back({1, 2, 3}); } // Repeat the same safety step for dataTwo { std::lock_guard<std::mutex> lock(mtxTwo); dataTwo.push_back({3, 4, 5}); } // Launch threads when we reach the 10000th element (i = 9999 since we start counting at 0) if (i == THRESHOLD - 1 && !threadsStarted) { // Pass references to vectors and mutexes using std::ref() std::thread t1(transferData, std::ref(dataOne), std::ref(dataThree), std::ref(mtxOne)); std::thread t2(transferData, std::ref(dataTwo), std::ref(dataFour), std::ref(mtxTwo)); // Wait for threads to finish before main thread exits (safer than detach) t1.detach(); // Optional: use detach if you don't need to wait, but join is more reliable t2.detach(); threadsStarted = true; std::cout << "Transfer threads launched! Main thread continues processing.\n"; } } // If using join instead of detach, uncomment this block to wait for threads: // if (threadsStarted) { // t1.join(); // t2.join(); // } // Verify the transfer worked std::cout << "Final dataThree size: " << dataThree.size() << "\n"; std::cout << "Final dataFour size: " << dataFour.size() << "\n"; return 0; }
Key Explanations
- Mutexes & Lock Guards:
std::mutexensures only one thread can access the vector at a time.std::lock_guardis a RAII wrapper that automatically unlocks the mutex when it goes out of scope—no manualunlock()calls needed, which reduces bugs. - Reference Passing: We use
std::ref()to pass references to our vectors and mutexes to the thread function. Without this, the thread would get copies of the data, which isn't what we want. - Thread Lifecycle: Using
join()ensures the main thread waits for transfer threads to finish before exiting.detach()lets threads run in the background, but you lose control over their completion—join()is safer for most use cases.
Bonus: Efficient Batch Processing
If you want to process data in batches (e.g., every 10000 elements, transfer that batch and start fresh), you can skip mutexes entirely by moving the vector to the thread instead of copying:
void batchTransfer(std::vector<std::vector<int>> src, std::vector<std::vector<int>>& dest) { // Move elements instead of copying (much faster for large datasets) dest.insert(dest.end(), std::make_move_iterator(src.begin()), std::make_move_iterator(src.end())); } // In main, replace the thread start logic with this: if (dataOne.size() >= THRESHOLD && !threadsStarted) { // Transfer ownership of dataOne to the thread (no copy overhead) std::thread t1(batchTransfer, std::move(dataOne), std::ref(dataThree)); dataOne.clear(); // Reset for new data std::thread t2(batchTransfer, std::move(dataTwo), std::ref(dataFour)); dataTwo.clear(); t1.join(); t2.join(); threadsStarted = true; }
This approach is more efficient because it avoids locking overhead and expensive data copies—perfect for your 10 million iteration use case.
内容的提问来源于stack exchange,提问作者Sajitha Liyanage

