You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

JavaScript搜索词匹配高亮及文本截断优化问题咨询

解决方案:多词匹配+连贯截断高亮

核心问题在于你把每个搜索词单独处理,导致生成多个独立片段。正确的做法是一次性处理所有搜索词,找到所有匹配位置后合并文本区间,再统一生成包含所有高亮词的连贯截断内容。

步骤1:重构核心处理函数

替换原来的searchText,实现多词匹配、区间合并、统一截断和高亮:

const processSearchDescription = (description, searchQuery) => {
  // 1. 移除HTML标签,转为纯文本(合并多余空格)
  const plainText = description.replace(/<\/?[^>]+(>|$)/g, " ").replace(/\s+/g, " ").trim();
  if (!plainText) return "";

  // 2. 处理搜索词:过滤空词,生成多词匹配正则
  const searchWords = searchQuery.trim().split(/\s+/).filter(word => word);
  if (searchWords.length === 0) return plainText.slice(0, 150) + "...";

  // 构建支持多词全局匹配的正则(忽略大小写)
  const regexPattern = searchWords.map(word => `(${word})`).join("|");
  const regex = new RegExp(regexPattern, "gi");

  // 3. 收集所有匹配的位置信息
  const matches = [];
  let match;
  while ((match = regex.exec(plainText)) !== null) {
    const start = match.index;
    const end = start + match[0].length;
    matches.push({ start, end });
    // 避免g模式下的无限循环(匹配空字符串时)
    if (match[0].length === 0) regex.lastIndex++;
  }
  if (matches.length === 0) return plainText.slice(0, 150) + "...";

  // 4. 合并重叠/邻近的匹配区间
  matches.sort((a, b) => a.start - b.start);
  const mergedIntervals = [matches[0]];
  for (let i = 1; i < matches.length; i++) {
    const last = mergedIntervals[mergedIntervals.length - 1];
    const current = matches[i];
    // 若当前区间和上一个重叠/距离小于radius,合并为一个区间
    if (current.start <= last.end + 30) {
      last.end = Math.max(last.end, current.end);
    } else {
      mergedIntervals.push(current);
    }
  }

  // 5. 确定最终截取范围:覆盖所有合并区间,前后各加radius
  const radius = 30;
  let overallStart = Math.min(...mergedIntervals.map(interval => interval.start)) - radius;
  overallStart = Math.max(overallStart, 0);
  let overallEnd = Math.max(...mergedIntervals.map(interval => interval.end)) + radius;
  overallEnd = Math.min(overallEnd, plainText.length);

  // 6. 添加省略号并高亮所有匹配词
  let truncatedText = plainText.slice(overallStart, overallEnd);
  if (overallStart > 0) truncatedText = "..." + truncatedText;
  if (overallEnd < plainText.length) truncatedText += "...";

  const highlightRegex = new RegExp(regexPattern, "gi");
  truncatedText = truncatedText.replace(highlightRegex, `<span class="font-bold text-yellow">$&</span>`);

  return truncatedText;
};

步骤2:重构组件渲染逻辑

不再循环每个搜索词生成片段,直接调用上面的函数生成单个连贯内容:

{
  searchResults.length > 0 ? (
    <div>
      <span className={`text-xs text-blue`}>Threads</span>
      {searchResults?.map((searched: any, index: Key) => {
        // 一次性处理整个搜索查询,生成包含所有高亮词的连贯文本
        const highlightedText = processSearchDescription(searched.description, search);

        return (
          <Link key={index} href={`/forum/thread/${searched?.slug}`}>
            <article
              className={`w-full bg-gray-dark mb-2 px-3 py-1.5 rounded-xl text-gray-light`}
            >
              <h2>{searched.title}</h2>
              <div
                className={`quill post-content text-xs text-gray`}
                dangerouslySetInnerHTML={{ __html: highlightedText }}
              />
            </article>
          </Link>
        );
      })}
    </div>
  ) : null;
}

关键优化点说明

  • 消除重复片段:通过合并所有匹配词的位置区间,只生成一段包含所有高亮词的连贯文本,避免每个词单独生成片段。
  • 多词匹配支持:用正则实现全局匹配所有搜索词,无需考虑词的顺序,所有匹配词都会被高亮。
  • 智能上下文保留:截取范围覆盖所有匹配词的周边内容,确保用户能看到搜索词的上下文信息。
  • 边界情况处理:包含了无匹配、空搜索词、空描述等场景的兼容处理。

内容的提问来源于stack exchange,提问作者Rade Iliev

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 09:46:06