Windows环境下用Python将UTF-16格式纯ASCII文件转单字节ASCII
Got it, let's solve this problem quickly and cleanly. Since your input file uses UTF-16 encoding but only contains standard ASCII characters, converting it to a single-byte ASCII file is straightforward with Python—no need to modify the original file at all.
Step-by-Step Solution
Here's a Python script tailored for Windows that handles the conversion safely:
# Set your input and output file paths (use raw strings to avoid Windows path escape issues) input_path = r"C:\example\input_file.txt" output_path = r"C:\example\output_file.txt" try: # Read the UTF-16 encoded file with open(input_path, 'r', encoding='utf-16') as input_file: file_content = input_file.read() # Write the content as standard ASCII with open(output_path, 'w', encoding='ascii') as output_file: output_file.write(file_content) print(f"Conversion complete! Output saved to: {output_path}") except FileNotFoundError: print(f"Error: Could not find input file at {input_path}") except Exception as e: print(f"Unexpected error during conversion: {str(e)}")
Key Details:
- Reading UTF-16: The
encoding='utf-16'parameter tells Python to correctly parse the double-byte UTF-16 format, including handling the Byte Order Mark (BOM) that's common in Windows UTF-16 files. - Writing ASCII: Since all characters are standard ASCII, using
encoding='ascii'will write each character as a single byte, exactly what you need for a non-double-byte ASCII file. - Windows Paths: Using raw strings (
r"") for file paths prevents issues with backslash escape characters in Windows file paths.
Edge Case Note
If for some reason there are hidden non-ASCII characters in the file (even though you mentioned all are ASCII), you can add an errors parameter to handle them. For example, encoding='ascii', errors='ignore' would skip non-ASCII characters, or errors='replace' would replace them with ?. But since you confirmed all characters are ASCII, this shouldn't be necessary.
内容的提问来源于stack exchange,提问作者Jeff Winchell

