You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Elixir中如何仅翻译字符串内指定列表中的短语而非整串?

Split Elixir Strings while Preserving Specific Phrases

Got it, let's tackle this problem where we need to split a string into elements that either match specific phrases from our list or are individual words. The key here is to prioritize longer phrases first (so we don't accidentally split a multi-word phrase into single words) and then extract both the target phrases and the remaining individual words correctly.

Approach 1: Using Regex Scan (Clean & Concise)

This method uses a regular expression that matches either our target phrases (sorted by length to prioritize longer ones) or any sequence of non-whitespace characters (individual words).

defmodule PhraseSplitter do
  def split_with_phrases(str, phrases) do
    # Sort phrases by length descending to prioritize longer matches
    sorted_phrases = phrases |> Enum.sort_by(&String.length/1, :desc)
    
    # Build regex pattern: match either a target phrase or a single word
    pattern = 
      sorted_phrases
      |> Enum.map(&Regex.escape/1)  # Escape special characters in phrases
      |> Enum.join("|")
      |> then(&~r/#{&1}|\S+/)
    
    # Scan the string for all matches
    Regex.scan(pattern, str) |> Enum.flat_map(& &1)
  end
end

# Example usage
list_with_words_that_needs_to_translate = ["here", "with one", "a string"]

IO.inspect PhraseSplitter.split_with_phrases("here is a string", list_with_words_that_needs_to_translate)
# Output: ["here", "is", "a string"]

IO.inspect PhraseSplitter.split_with_phrases("word with one and two", list_with_words_that_needs_to_translate)
# Output: ["word", "with one", "and two"]

IO.inspect PhraseSplitter.split_with_phrases("and something that doesn't need translate", list_with_words_that_needs_to_translate)
# Output: ["and", "something", "that", "doesn't", "need", "translate"]

How This Works:

  1. Sorting Phrases: By sorting phrases from longest to shortest, we ensure that multi-word phrases like "with one" are matched before any shorter overlapping phrases (if they existed).
  2. Regex Pattern: The pattern combines escaped versions of our phrases (to handle any special characters safely) with a fallback to match single words (\S+).
  3. Regex Scan: Regex.scan finds all non-overlapping matches in the string, which directly gives us our desired list of elements.

Approach 2: Using String.Split (More Explicit)

If you prefer a more step-by-step approach, you can split the string using the target phrases as delimiters, then process the remaining parts to split into individual words:

defmodule PhraseSplitter do
  def split_with_phrases(str, phrases) do
    sorted_phrases = phrases |> Enum.sort_by(&String.length/1, :desc)
    
    pattern = 
      sorted_phrases
      |> Enum.map(&Regex.escape/1)
      |> Enum.join("|")
      |> then(&~r/(#{&1})/)  # Capture phrases to keep them in split results
    
    str
    |> String.split(pattern, trim: true, include_captures: true)
    |> Enum.flat_map(fn part ->
      if part in phrases do
        [part]  # Keep the matched phrase as a single element
      else
        # Split non-phrase parts into individual words, skipping empty strings
        part |> String.split() |> Enum.reject(&(&1 == ""))
      end
    end)
  end
end

This works similarly, but uses String.split with include_captures: true to retain the target phrases, then processes each split part to either keep the phrase or split into words.

Both approaches will handle edge cases like phrases at the start/end of the string, multiple consecutive spaces, and overlapping phrase scenarios correctly.

内容的提问来源于stack exchange,提问作者Elmseld

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 12:22:28