You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用JavaScript拆分文本单词并存储每个单词的起止索引?

Solution to Extract Words with Start/End Indices

Got it, I’ve been stuck on this exact problem before—grabbing words from text is straightforward with match(), but getting their position indices is where things get annoying. No need to write a full custom parser though; here’s a clean, efficient approach using JavaScript’s regex exec() method:

Step-by-Step Code Implementation

const inputText = "Hello world! This is a test string with words to index.";
const wordRegex = /\b(\w+)\b/g;
const wordIndexArray = [];
let currentMatch;

// Loop through all matches using exec()
while ((currentMatch = wordRegex.exec(inputText)) !== null) {
  const word = currentMatch[1];
  const startIndex = currentMatch.index;
  const endIndex = startIndex + word.length;
  wordIndexArray.push([word, startIndex, endIndex]);
}

console.log(wordIndexArray);
// Output example: [["Hello",0,5], ["world",6,11], ["This",13,17], ...]

How This Works

  • The exec() method doesn’t just return matches—it gives a full match object that includes index, the starting position of the matched word in the original text.
  • We calculate the endIndex by adding the word’s length to the start index (since strings are zero-indexed in JS).
  • The while loop keeps running until exec() returns null, meaning no more matches are found.

Customization Tip

If your text includes words with apostrophes (like don’t or Mike’s) or hyphenated terms, adjust the regex to include those cases:

// Matches words with optional apostrophes
const flexibleWordRegex = /\b(\w+(?:['’]\w+)?)\b/g;

This should cover your use case perfectly, and it’s way simpler than building a full text parser from scratch.

内容的提问来源于stack exchange,提问作者Herman Neple

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 11:40:56