Qt应用中UTF-7解码生成错误字符问题排查求助
UTF-7解码字符错误问题排查(Qt文本编辑器)
开发基于Qt的文本编辑器时,解码UTF-7编码文本遇到字符错误:对UTF-7字符串Hello, +Z1TA-(对应中文Hello, 世界)解码,结果得到Hello, ᙎ䱵。已实现InterpreteAsUtf7类的解码逻辑,疑似Base64段处理或UTF-16转换环节存在问题,以下是复现代码:
#include <QString> #include <QByteArray> #include <QDebug> class InterpreteAsUtf7 { public: static QString decodeUtf7(const QByteArray& utf7Data); private: static QString processBase64Segment(const QByteArray& base64Buffer, bool& errorFlag); static QString convertToUtf16LE(const QByteArray& decodedBytes, bool& errorFlag); }; QString InterpreteAsUtf7::decodeUtf7(const QByteArray& utf7Data) { QString result; QByteArray base64Buffer; bool inBase64 = false; bool decodingError = false; for (char c : utf7Data) { if (c == '+') { if (inBase64) { result += "+" + base64Buffer; // Unclosed Base64 section treated as literal base64Buffer.clear(); } inBase64 = true; continue; } if (c == '-' && inBase64) { result += processBase64Segment(base64Buffer, decodingError); base64Buffer.clear(); inBase64 = false; continue; } if (inBase64) { base64Buffer.append(c); } else { result += QChar(c); // Literal character } } if (inBase64) { result += "+" + base64Buffer; // Unclosed Base64 treated as literal } return result; } QString InterpreteAsUtf7::processBase64Segment(const QByteArray& base64Buffer, bool& errorFlag) { QByteArray decodedBytes = QByteArray::fromBase64(base64Buffer); if (decodedBytes.isEmpty()) { errorFlag = true; qWarning() << "[WARNING] Base64 decoding failed for:" << base64Buffer; return QString(); } return convertToUtf16LE(decodedBytes, errorFlag); } QString InterpreteAsUtf7::convertToUtf16LE(const QByteArray& decodedBytes, bool& errorFlag) { if (decodedBytes.size() % 2 != 0) { errorFlag = true; qWarning() << "[WARNING] Decoded bytes not aligned for UTF-16:" << decodedBytes; return QString(); } QString result; for (int i = 0; i < decodedBytes.size(); i += 2) { ushort codeUnit = static_cast<uchar>(decodedBytes[i]) | (static_cast<uchar>(decodedBytes[i + 1]) << 8); result.append(QChar(codeUnit)); } return result; } int main() { QByteArray utf7Data = "Hello, +Z1TA-"; // UTF-7 encoded string for "Hello, 世界" QString decodedText = InterpreteAsUtf7::decodeUtf7(utf7Data); qDebug() << "Original UTF-7:" << utf7Data; qDebug() << "Decoded UTF-7:" << decodedText; return 0; }
错误原因分析
- UTF-16字节顺序错误:UTF-7标准中,Base64解码后的字节是**UTF-16大端(BE)**编码,但代码中按小端(LE)拼接字节,导致字符编码反转。
- Base64格式不匹配:UTF-7使用修改版Base64,将标准Base64的
/替换为,,且不需要填充字符=。直接调用QByteArray::fromBase64会因格式不兼容导致解码错误。
修正后的代码
修改processBase64Segment和convertToUtf16LE函数(重命名为convertToUtf16BE更准确):
QString InterpreteAsUtf7::processBase64Segment(const QByteArray& base64Buffer, bool& errorFlag) { // 替换UTF-7修改版Base64的字符为标准Base64格式 QByteArray adjustedBase64 = base64Buffer; adjustedBase64.replace(',', '/'); // UTF-7不需要填充,手动补充必要的'='使长度为4的倍数 int padding = adjustedBase64.size() % 4; if (padding != 0) { adjustedBase64.append(4 - padding, '='); } QByteArray decodedBytes = QByteArray::fromBase64(adjustedBase64); if (decodedBytes.isEmpty()) { errorFlag = true; qWarning() << "[WARNING] Base64 decoding failed for:" << base64Buffer; return QString(); } return convertToUtf16BE(decodedBytes, errorFlag); } QString InterpreteAsUtf7::convertToUtf16BE(const QByteArray& decodedBytes, bool& errorFlag) { if (decodedBytes.size() % 2 != 0) { errorFlag = true; qWarning() << "[WARNING] Decoded bytes not aligned for UTF-16:" << decodedBytes; return QString(); } QString result; for (int i = 0; i < decodedBytes.size(); i += 2) { // 按大端顺序拼接字节:高字节在前,低字节在后 ushort codeUnit = (static_cast<uchar>(decodedBytes[i]) << 8) | static_cast<uchar>(decodedBytes[i + 1]); result.append(QChar(codeUnit)); } return result; }
修正后效果
运行修正后的代码,Hello, +Z1TA-将正确解码为Hello, 世界。
内容的提问来源于stack exchange,提问作者Remisa Phillips
相关产品推荐
相关产品推荐

