如何修改JavaScript代码实现双词分组而非单词分组
Got it, let's get this sorted! You're right that the regex is the main piece to adjust here—your original code captures single alphabetic words, but we need to modify it to grab pairs of words separated by whitespace. Here's how to do it step by step:
Key Changes Needed
- Update the regular expression to match two consecutive alphabetic words separated by one or more spaces.
- Adjust the logic to combine the two captured words into a single "pair" key for your result object.
- Store the full range of the word pair (start index to end index of the entire pair) instead of individual word positions.
Modified Code
const countWordPairs = (text) => { // Regex to capture two alphabetic words separated by one or more spaces const wordPairRegex = new RegExp(`([\\p{Alphabetic}]+)\\s+([\\p{Alphabetic}]+)`, 'gu'); let result; const wordPairs = {}; while ((result = wordPairRegex.exec(text)) !== null) { // Combine the two words into a lowercase pair (e.g., "hello my") const pair = `${result[1].toLowerCase()} ${result[2].toLowerCase()}`; // Get the full start and end indices of the entire word pair const startIndex = result.index; const endIndex = result.index + result[0].length; if (!wordPairs[pair]) { wordPairs[pair] = []; } // Store the start and end positions of this pair occurrence wordPairs[pair].push(startIndex, endIndex); } return wordPairs; };
How It Works
Regex Breakdown:
([\p{Alphabetic}]+)\s+([\p{Alphabetic}]+)([\p{Alphabetic}]+): Captures the first word (any alphabetic characters, supports Unicode)\s+: Matches one or more spaces between the two words([\p{Alphabetic}]+): Captures the second word- The
guflags ensure we match all pairs globally and handle Unicode correctly.
Pair Handling: We combine the two captured words into a lowercase string (like
"hello my") to use as the key in ourwordPairsobject—this ensures case-insensitive matching (e.g., "Hello My" and "hello my" are treated as the same pair).Index Storage: Instead of storing individual word indices, we store the start index of the entire pair and the end index (calculated as start index plus the length of the matched pair string).
Edge Case Note
If your input text has an odd number of words (e.g., "hello my friend"), the last single word won't be captured by this regex. If you need to handle that, you could add an optional check for remaining single words, but based on your original request, focusing on pairs should cover your use case.
内容的提问来源于stack exchange,提问作者Cevin Thomas

