Ubuntu下十进制转Unicode转换器文件输出失败问题求助
Hey there! Let's figure out how to get your Unicode converter writing to files properly—console output working but files failing is such a common gotcha with Unicode handling, so I’ve got you covered.
First, let's break down the core issue
Your extended fstream works for the console because it’s handling multi-byte Unicode (like the 3-byte UTF-8 sequence for €) correctly in that context. For files, though, we need to ensure we’re writing the right byte sequence or configuring the stream to handle Unicode encoding properly.
Option 1: Write raw UTF-8 bytes directly to file
If your converter already turns decimal numbers into valid UTF-8 byte sequences (like 226 130 172 for €), this is the simplest approach. We’ll use a standard ofstream in binary mode to avoid corrupting the multi-byte data with automatic newline conversions.
#include <fstream> #include <vector> // Example: Your converter outputs this UTF-8 byte sequence for € std::vector<unsigned char> utf8_bytes = {226, 130, 172}; // Open file in binary mode to preserve raw bytes std::ofstream output_file("result.txt", std::ios::binary); if (output_file.is_open()) { // Write the raw bytes to the file output_file.write(reinterpret_cast<const char*>(utf8_bytes.data()), utf8_bytes.size()); output_file.close(); }
When you open result.txt with a text editor that supports UTF-8, you’ll see the correct € symbol.
Option 2: Use wide character streams (for your extended fstream)
If your extended fstream relies on wide characters (wchar_t), you need to configure the stream to encode wide characters into UTF-8 before writing to the file. The default locale won’t do this automatically—we have to explicitly set a UTF-8-aware codecvt facet.
#include <fstream> #include <locale> #include <codecvt> #include <string> // Convert your wide character string to UTF-8 std::wstring_convert<std::codecvt_utf8<wchar_t>> utf8_converter; std::wstring unicode_wstr = L"€"; // Example wide char from your converter std::string utf8_str = utf8_converter.to_bytes(unicode_wstr); // Write the UTF-8 string to file std::ofstream output_file("result.txt"); if (output_file.is_open()) { output_file << utf8_str; output_file.close(); }
Alternatively, you can imbue the wide stream directly with a UTF-8 locale:
std::wofstream output_file("result.txt"); // Set locale to handle UTF-8 encoding for wide characters output_file.imbue(std::locale(output_file.getloc(), new std::codecvt_utf8<wchar_t>)); if (output_file.is_open()) { output_file << L"€"; // Write your wide character directly output_file.close(); }
Note: codecvt_utf8 is marked deprecated in C17, but it’s still fully supported in GCC, Clang, and MSVC. For C20+, you could use <format> or newer Unicode utilities, but the above code will work for most use cases.
Key things to remember
- Stick to one encoding: Always write files in a consistent encoding (UTF-8 is the safest choice) and make sure you open the file with that encoding in your text editor.
- Avoid text mode for raw bytes: When writing multi-byte sequences, use
std::ios::binaryto prevent the OS from altering newline characters, which can break Unicode sequences. - Update your extended fstream: If you want to integrate this into your custom stream class, add a method that either writes raw UTF-8 bytes (like Option 1) or converts wide chars to UTF-8 before writing.
内容的提问来源于stack exchange,提问作者user9463791

