如何获取文件中指定行列位置的单词?
How to Extract a Word from a Specific Line and Column in a File with Ruby
Got it, let's work through this problem together. You need to build a get_word function that finds the word containing the character at a given line and column position in a file. Here's how to fill in the missing part of your code, with explanations for each step:
Step-by-Step Explanation
Once you've located the target line, the key is to find the boundaries of the word that includes the specified column:
- Convert the 1-based column number to 0-based (since Ruby strings use 0-indexing).
- Handle edge cases where the column is out of bounds (e.g., column number is larger than the line length).
- Find the left boundary of the word: move left from the target character until you hit a non-letter character.
- Find the right boundary: move right until you hit a non-letter character.
- Extract and return the substring between these boundaries.
We'll use Unicode-aware letter matching (\p{Letter}) to handle accented characters like Á or ñ correctly, and specify UTF-8 encoding to avoid issues with non-ASCII text.
Complete Code
def get_word(file, position) target_line_num = position.line_number target_col = position.column_number # Iterate through each line with UTF-8 encoding to handle accented characters File.each_line(file, encoding: 'UTF-8').with_index do |line, line_idx| current_line = line_idx + 1 # Convert 0-based index to 1-based line number next unless current_line == target_line_num # Remove trailing newline so it doesn't count as a character in the column count line_chars = line.chomp.chars char_index = target_col - 1 # Convert 1-based column to 0-based index # Return nil if column is out of bounds return nil if char_index < 0 || char_index >= line_chars.length # Find left boundary of the word (first letter of the word) left_boundary = char_index while left_boundary >= 0 && line_chars[left_boundary] =~ /\p{Letter}/ left_boundary -= 1 end left_boundary += 1 # Find right boundary of the word (last letter of the word) right_boundary = char_index while right_boundary < line_chars.length && line_chars[right_boundary] =~ /\p{Letter}/ right_boundary += 1 end right_boundary -= 1 # Join the characters between boundaries to get the word return line_chars[left_boundary..right_boundary].join end # Return nil if the target line doesn't exist in the file nil end
Testing with Your Example
For your message.md file:
Dear people: My name is [Ángeliño](#angelino). Bye!
get_word("message.md", Position.new(1, 9))returns"people"(the 9th column is the second 'p' in "people", and the function correctly captures the entire word).get_word("message.md", Position.new(2, 13))returns"Ángeliño"(the 13th column is theÁ, and the function handles the accented characters properly).
Key Notes
- Unicode Support: Using
\p{Letter}ensures we match all alphabetic characters across languages, not just ASCII letters. - Encoding: Specifying
encoding: 'UTF-8'when reading the file prevents garbled text for non-ASCII characters. - Edge Cases: The function returns
nilif the target line doesn't exist or the column is out of bounds, making it robust against invalid inputs.
内容的提问来源于stack exchange,提问作者user7450793
相关产品推荐
相关产品推荐

