如何用JavaScript拆分文本单词并存储每个单词的起止索引?
Solution to Extract Words with Start/End Indices
Got it, I’ve been stuck on this exact problem before—grabbing words from text is straightforward with match(), but getting their position indices is where things get annoying. No need to write a full custom parser though; here’s a clean, efficient approach using JavaScript’s regex exec() method:
Step-by-Step Code Implementation
const inputText = "Hello world! This is a test string with words to index."; const wordRegex = /\b(\w+)\b/g; const wordIndexArray = []; let currentMatch; // Loop through all matches using exec() while ((currentMatch = wordRegex.exec(inputText)) !== null) { const word = currentMatch[1]; const startIndex = currentMatch.index; const endIndex = startIndex + word.length; wordIndexArray.push([word, startIndex, endIndex]); } console.log(wordIndexArray); // Output example: [["Hello",0,5], ["world",6,11], ["This",13,17], ...]
How This Works
- The
exec()method doesn’t just return matches—it gives a full match object that includesindex, the starting position of the matched word in the original text. - We calculate the
endIndexby adding the word’s length to the start index (since strings are zero-indexed in JS). - The
whileloop keeps running untilexec()returnsnull, meaning no more matches are found.
Customization Tip
If your text includes words with apostrophes (like don’t or Mike’s) or hyphenated terms, adjust the regex to include those cases:
// Matches words with optional apostrophes const flexibleWordRegex = /\b(\w+(?:['’]\w+)?)\b/g;
This should cover your use case perfectly, and it’s way simpler than building a full text parser from scratch.
内容的提问来源于stack exchange,提问作者Herman Neple
相关产品推荐
相关产品推荐

