如何使用Ruby从非UTF-8文本/Typescript文件提取数字并写入新文件?
Hey there! Let's tackle your two Ruby file processing challenges head-on—both are about extracting numbers, but their root issues are distinct. Here's how to fix each scenario:
1. Extracting Numbers from Non-UTF-8 Encoded Files
First, you need to make sure you're reading the file with its actual encoding (otherwise you'll get garbled text that breaks your number extraction).
Step 1: Identify the file's encoding
Use your terminal to run:
file --mime-encoding your_input_file.txt
This will tell you the real encoding (e.g., ISO-8859-1, GBK, Shift_JIS).
Step 2: Ruby code to extract numbers
Use the identified encoding to read the file, then scan for numeric sequences and write them to a new UTF-8 file for broad compatibility:
# Replace with your file's actual encoding from the terminal command source_encoding = "ISO-8859-1" input_path = "non_utf8_source.txt" output_path = "extracted_numbers.txt" # Write output as UTF-8 for broad compatibility File.open(output_path, "w:UTF-8") do |output_file| # Read the source file with its correct encoding File.open(input_path, "r:#{source_encoding}") do |input_file| input_file.each_line do |line| # Extract integers and decimals (adjust regex if you need other number formats) numbers = line.scan(/\d+(\.\d+)?/) output_file.puts numbers.flatten.join("\n") unless numbers.empty? end end end
2. Fixing "Broken UTF-8" Files (e.g., Converted Typescript Logs)
Your issue here is likely invisible control characters (like terminal ANSI escape sequences) or invalid UTF-8 bytes that slip through encoding checks. Standard gsub!/tr fails because these characters aren't matched by basic regex patterns.
Solution 1: Clean invalid bytes and terminal codes first
Read the file as binary, force it to valid UTF-8, then strip terminal control codes before extracting numbers:
input_path = "broken_typescript_converted.txt" output_path = "clean_extracted_numbers.txt" File.open(output_path, "w:UTF-8") do |output_file| # Read as binary to capture all bytes, then fix invalid UTF-8 raw_content = File.binread(input_path) cleaned_content = raw_content.encode("UTF-8", invalid: :replace, undef: :replace, replace: "") # Remove terminal control codes (common in typescript logs) cleaned_content.gsub!(/\e\[[0-9;]*[a-zA-Z]/, "") # Extract all valid numbers numbers = cleaned_content.scan(/\d+(\.\d+)?/) output_file.puts numbers.flatten.join("\n") end
Solution 2: Strictly filter for digits only
If the above doesn't work, directly strip everything except digits and optional decimals to eliminate any hidden junk:
input_path = "broken_typescript_converted.txt" output_path = "clean_extracted_numbers.txt" File.open(output_path, "w:UTF-8") do |output_file| File.open(input_path, "r:UTF-8") do |input_file| input_file.each_line do |line| # Keep only digits and dots, then clean up invalid sequences (like multiple dots) stripped_line = line.gsub(/[^\d.]/, "") valid_numbers = stripped_line.split(/(?<=\d)\.(?=\d)/).reject { |s| s.empty? } output_file.puts valid_numbers unless valid_numbers.empty? end end end
This ensures even weird invisible characters get stripped out before you extract your target numbers.
内容的提问来源于stack exchange,提问作者sharknado

