You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Ruby从非UTF-8文本/Typescript文件提取数字并写入新文件?

Hey there! Let's tackle your two Ruby file processing challenges head-on—both are about extracting numbers, but their root issues are distinct. Here's how to fix each scenario:

1. Extracting Numbers from Non-UTF-8 Encoded Files

First, you need to make sure you're reading the file with its actual encoding (otherwise you'll get garbled text that breaks your number extraction).

Step 1: Identify the file's encoding

Use your terminal to run:

file --mime-encoding your_input_file.txt

This will tell you the real encoding (e.g., ISO-8859-1, GBK, Shift_JIS).

Step 2: Ruby code to extract numbers

Use the identified encoding to read the file, then scan for numeric sequences and write them to a new UTF-8 file for broad compatibility:

# Replace with your file's actual encoding from the terminal command
source_encoding = "ISO-8859-1"
input_path = "non_utf8_source.txt"
output_path = "extracted_numbers.txt"

# Write output as UTF-8 for broad compatibility
File.open(output_path, "w:UTF-8") do |output_file|
  # Read the source file with its correct encoding
  File.open(input_path, "r:#{source_encoding}") do |input_file|
    input_file.each_line do |line|
      # Extract integers and decimals (adjust regex if you need other number formats)
      numbers = line.scan(/\d+(\.\d+)?/)
      output_file.puts numbers.flatten.join("\n") unless numbers.empty?
    end
  end
end

2. Fixing "Broken UTF-8" Files (e.g., Converted Typescript Logs)

Your issue here is likely invisible control characters (like terminal ANSI escape sequences) or invalid UTF-8 bytes that slip through encoding checks. Standard gsub!/tr fails because these characters aren't matched by basic regex patterns.

Solution 1: Clean invalid bytes and terminal codes first

Read the file as binary, force it to valid UTF-8, then strip terminal control codes before extracting numbers:

input_path = "broken_typescript_converted.txt"
output_path = "clean_extracted_numbers.txt"

File.open(output_path, "w:UTF-8") do |output_file|
  # Read as binary to capture all bytes, then fix invalid UTF-8
  raw_content = File.binread(input_path)
  cleaned_content = raw_content.encode("UTF-8", invalid: :replace, undef: :replace, replace: "")
  
  # Remove terminal control codes (common in typescript logs)
  cleaned_content.gsub!(/\e\[[0-9;]*[a-zA-Z]/, "")
  
  # Extract all valid numbers
  numbers = cleaned_content.scan(/\d+(\.\d+)?/)
  output_file.puts numbers.flatten.join("\n")
end

Solution 2: Strictly filter for digits only

If the above doesn't work, directly strip everything except digits and optional decimals to eliminate any hidden junk:

input_path = "broken_typescript_converted.txt"
output_path = "clean_extracted_numbers.txt"

File.open(output_path, "w:UTF-8") do |output_file|
  File.open(input_path, "r:UTF-8") do |input_file|
    input_file.each_line do |line|
      # Keep only digits and dots, then clean up invalid sequences (like multiple dots)
      stripped_line = line.gsub(/[^\d.]/, "")
      valid_numbers = stripped_line.split(/(?<=\d)\.(?=\d)/).reject { |s| s.empty? }
      
      output_file.puts valid_numbers unless valid_numbers.empty?
    end
  end
end

This ensures even weird invisible characters get stripped out before you extract your target numbers.

内容的提问来源于stack exchange,提问作者sharknado

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 09:15:56