Windows下如何从其他进程内存读取Unicode字符串?
Let's break down what's going wrong here and fix it step by step—your code has a critical memory read error plus an encoding issue that's blocking Cyrillic characters from working correctly.
1. Fix the Broken Memory Read Buffer
First, the biggest mistake is in your ReadWideString function: you declared outStr as a temporary buffer for the string, but you're passing &out (the address of the std::wstring object itself) to ReadProcessMemory. That's completely wrong—std::wstring stores metadata (like pointers and length) internally, not the actual character data. Reading memory directly into this object corrupts its structure, so you're never actually capturing the target string.
Here's the corrected function, with proper buffer handling and safety checks:
// Define a reasonable maximum string length first (adjust as needed) constexpr size_t maxStringLength = 1024; bool ReadWideString(const HANDLE& hProc, const std::uintptr_t& addr, std::wstring& out) { std::array<wchar_t, maxStringLength> outStr{}; // Initialize to avoid garbage data SIZE_T bytesRead = 0; // Read the wide string from the target process into our temporary buffer auto readMemRes = ReadProcessMemory( hProc, reinterpret_cast<LPCVOID>(addr), outStr.data(), outStr.size() * sizeof(wchar_t), // Calculate total bytes (wchar_t is 2 bytes on Windows) &bytesRead ); if (!readMemRes || bytesRead == 0) { return false; } // Find the null terminator to avoid including garbage from the array size_t actualLength = 0; while (actualLength < outStr.size() && outStr[actualLength] != L'\0') { ++actualLength; } // Assign only the valid character data to the output string out.assign(outStr.data(), actualLength); return true; }
2. Fix Cyrillic Character Output
Even with a correct read, Windows' std::wofstream uses the system's local code page by default (like CP1251), which can mangle Cyrillic characters. You have two reliable fixes:
Option A: Convert to UTF-8 (Recommended for Compatibility)
Convert the UTF-16 std::wstring to a UTF-8 std::string, then write with a regular std::ofstream:
#include <codecvt> // Required for string conversion // After calling ReadWideString std::wstring name; bool res = ReadWideString(OpenedProcessHandle, address, name); if (res) { // Convert UTF-16 wstring to UTF-8 string std::wstring_convert<std::codecvt_utf8_utf16<wchar_t>> utfConverter; std::string utf8Name = utfConverter.to_bytes(name); std::ofstream test("test.txt"); test << utf8Name; test.close(); }
Option B: Write Directly as UTF-16
If you want to keep the UTF-16 encoding, set a proper locale for wofstream:
#include <locale> // Required for locale configuration // After calling ReadWideString std::wstring name; bool res = ReadWideString(OpenedProcessHandle, address, name); if (res) { std::wofstream test("test.txt"); // Set locale to UTF-16 little-endian (Windows' native Unicode format) test.imbue(std::locale( test.getloc(), new std::codecvt_utf16<wchar_t, 0x10ffff, std::little_endian> )); test << name; test.close(); }
3. Why std::string Gave You Garbage
On Windows, wchar_t is 2 bytes. Cyrillic characters like Д use the UTF-16 value 0x0414, which is stored in memory as 0x14 0x04 (little-endian). When you treat this 2-byte data as a std::string, each byte is interpreted as a separate ASCII character—those values map to unprintable or random characters, hence the "EC9"-like garbage.
内容的提问来源于stack exchange,提问作者Shiroe

