Elixir中如何仅翻译字符串内指定列表中的短语而非整串?
Got it, let's tackle this problem where we need to split a string into elements that either match specific phrases from our list or are individual words. The key here is to prioritize longer phrases first (so we don't accidentally split a multi-word phrase into single words) and then extract both the target phrases and the remaining individual words correctly.
Approach 1: Using Regex Scan (Clean & Concise)
This method uses a regular expression that matches either our target phrases (sorted by length to prioritize longer ones) or any sequence of non-whitespace characters (individual words).
defmodule PhraseSplitter do def split_with_phrases(str, phrases) do # Sort phrases by length descending to prioritize longer matches sorted_phrases = phrases |> Enum.sort_by(&String.length/1, :desc) # Build regex pattern: match either a target phrase or a single word pattern = sorted_phrases |> Enum.map(&Regex.escape/1) # Escape special characters in phrases |> Enum.join("|") |> then(&~r/#{&1}|\S+/) # Scan the string for all matches Regex.scan(pattern, str) |> Enum.flat_map(& &1) end end # Example usage list_with_words_that_needs_to_translate = ["here", "with one", "a string"] IO.inspect PhraseSplitter.split_with_phrases("here is a string", list_with_words_that_needs_to_translate) # Output: ["here", "is", "a string"] IO.inspect PhraseSplitter.split_with_phrases("word with one and two", list_with_words_that_needs_to_translate) # Output: ["word", "with one", "and two"] IO.inspect PhraseSplitter.split_with_phrases("and something that doesn't need translate", list_with_words_that_needs_to_translate) # Output: ["and", "something", "that", "doesn't", "need", "translate"]
How This Works:
- Sorting Phrases: By sorting phrases from longest to shortest, we ensure that multi-word phrases like
"with one"are matched before any shorter overlapping phrases (if they existed). - Regex Pattern: The pattern combines escaped versions of our phrases (to handle any special characters safely) with a fallback to match single words (
\S+). - Regex Scan:
Regex.scanfinds all non-overlapping matches in the string, which directly gives us our desired list of elements.
Approach 2: Using String.Split (More Explicit)
If you prefer a more step-by-step approach, you can split the string using the target phrases as delimiters, then process the remaining parts to split into individual words:
defmodule PhraseSplitter do def split_with_phrases(str, phrases) do sorted_phrases = phrases |> Enum.sort_by(&String.length/1, :desc) pattern = sorted_phrases |> Enum.map(&Regex.escape/1) |> Enum.join("|") |> then(&~r/(#{&1})/) # Capture phrases to keep them in split results str |> String.split(pattern, trim: true, include_captures: true) |> Enum.flat_map(fn part -> if part in phrases do [part] # Keep the matched phrase as a single element else # Split non-phrase parts into individual words, skipping empty strings part |> String.split() |> Enum.reject(&(&1 == "")) end end) end end
This works similarly, but uses String.split with include_captures: true to retain the target phrases, then processes each split part to either keep the phrase or split into words.
Both approaches will handle edge cases like phrases at the start/end of the string, multiple consecutive spaces, and overlapping phrase scenarios correctly.
内容的提问来源于stack exchange,提问作者Elmseld

