如何拼接两个字符串并省略首尾重复的连续语句段?
合并两段存在末尾-开头重复的文本
问题场景
开发语音转文本模块时,需要合并两段文本:从第一段文本的末尾向前查找,找到与第二段文本开头重复的语句段,最终合并为无重复的完整语句。例如:
const str1 = "It was great, I finally had some time to" const str2 = "had some time to relax and catch up on my reading." // 期望结果:"It was great, I finally had some time to relax and catch up on my reading."
原代码的问题
你提供的formPhrase函数通过去重所有单词来实现合并,这种逻辑只适用于重复内容仅出现在str1末尾和str2开头的场景。当重复内容同时存在于文本的其他位置时就会失效,比如:
const str3 = "this code transforms this"; const str4 = "transforms this string"; // 原函数输出:"this code transforms string" // 正确期望输出:"this code transforms this string"
原函数会错误地移除str1中中间出现的重复单词,不符合“仅合并末尾-开头重复”的需求。
正确实现方案
我们需要找到str1后缀与str2前缀的最长匹配片段,再进行合并,确保只去除末尾和开头的重复部分:
function mergeWithOverlap(str1, str2) { const words1 = str1.split(' '); const words2 = str2.split(' '); const maxPossibleOverlap = Math.min(words1.length, words2.length); // 从最长的可能重叠长度开始检查,确保找到的是最长匹配 for (let overlapLength = maxPossibleOverlap; overlapLength > 0; overlapLength--) { const str1Suffix = words1.slice(-overlapLength).join(' '); const str2Prefix = words2.slice(0, overlapLength).join(' '); if (str1Suffix === str2Prefix) { // 合并:str1完整内容 + str2去掉重叠前缀的部分 return [...words1, ...words2.slice(overlapLength)].join(' '); } } // 无重叠时直接拼接两段文本 return `${str1} ${str2}`; } // 测试示例1 const str1 = "It was great, I finally had some time to"; const str2 = "had some time to relax and catch up on my reading."; console.log(mergeWithOverlap(str1, str2)); // 输出:It was great, I finally had some time to relax and catch up on my reading. // 测试示例2 const str3 = "this code transforms this"; const str4 = "transforms this string"; console.log(mergeWithOverlap(str3, str4)); // 输出:this code transforms this string
逻辑说明
- 将两段文本拆分为单词数组,便于按片段检查;
- 从最长的可能重叠长度(两段文本单词数的较小值)开始遍历,确保找到的是最长的匹配重叠段;
- 找到匹配的后缀和前缀后,合并两段文本(str1完整保留,str2去掉重叠的前缀部分);
- 若没有任何重叠,直接拼接两段文本。
内容的提问来源于stack exchange,提问作者Maurício Giordano
相关产品推荐
相关产品推荐

