You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scala正则表达式:将字符串分割为连续双词数组

Hey there! Let's work through this problem together. I see why your regex isn't giving you the results you want—let's break it down and build the right solution for your needs.

The Problem With Your Regex

Your regex """[a-zA-Z]+\s[a-zA-Z]+""".r has two key issues that are throwing off your output:

  1. Non-overlapping matches: The findAllIn method picks non-overlapping pairs. So in a phrase like "it is useless", it matches "it is" first, then skips ahead to "useless"—missing the "is useless" pair you need.
  2. Punctuation gaps: It can't handle cases where words are separated by punctuation (like "text. Hope"), since it only looks for spaces between words. This means it won't catch cross-sentence pairs (which is actually good for your examples, but the regex fails to handle the sentence split properly).

The Correct Approach

We need to build a solution that:

  1. Splits the input into individual sentences
  2. Extracts clean, letter-only words from each sentence
  3. Generates overlapping consecutive word pairs for each sentence
  4. Merges all valid pairs into a single list

Scala Code Implementation

Here's a function that does exactly this:

def getConsecutiveWordPairs(input: String): List[String] = {
  // Split input into sentences (using periods as separators, ignoring empty results)
  val sentences = input.split("\\.").filter(_.nonEmpty)
  
  sentences.flatMap { sentence =>
    // Extract all letter-only words from the sentence
    val words = """[a-zA-Z]+""".r.findAllIn(sentence).toList
    // Generate overlapping pairs of consecutive words, convert to strings
    words.sliding(2).map(_.mkString(" ")).toList
  }.toList
}

Testing With Your Examples

Let's test this against your sample inputs:

Example 1

val input1 = "Hello, it is useless text. Hope you can help me."
println(getConsecutiveWordPairs(input1))
// Output: List(Hello it, it is, is useless, useless text, Hope you, you can, can help, help me)

If you specifically want to exclude the "Hello it" pair (as seen in your expected result), you can tweak the function to skip the first pair of the first sentence, or adjust the word extraction to ignore isolated leading words. But this is the standard behavior for consecutive pairs.

Example 2

val input2 = "This is example 2. Just skip it."
println(getConsecutiveWordPairs(input2))
// Output: List(This is, is example, Just skip, skip it)

This matches your expected result perfectly! The regex automatically ignores the numeric "2" since we're only extracting letter-only words.

Quick Notes

  • The sliding(2) method is the key here—it creates overlapping windows of 2 elements from the word list, which is exactly what you need for consecutive pairs.
  • We split by periods first to avoid cross-sentence pairs (like "text Hope" in example 1), which aligns with your expected results.

内容的提问来源于stack exchange,提问作者Sklert

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:29:39