You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何利用哈希表优化Gmail中千级主题文本匹配的性能

Fixing Gmail Performance Issues with 1000+ Topic Matches

Let's break down why your current code is causing Gmail to choke, then fix it with a hash map (and smart regex optimization) to get things running smoothly.

The Root of the Problem

Your current approach runs searchPage() once per topic—that's 1000+ separate passes over Gmail's massive DOM tree. Each pass uses a TreeWalker and regex, which adds up to way too much work for the browser. No wonder it's crashing!

The Fix: One DOM Pass, All Matches

Instead of looping through each topic and re-scanning the DOM every time, we'll:

  1. Combine all topics into a single regex (so we can match every topic in one go)
  2. Use a Map (our hash table) to quickly look up topic data when we find a match
  3. Only traverse the DOM once to find and highlight all matches

Here's the optimized code:

var topicData = [ 
  { id: 1, name: 'Courses' }, 
  { id: 2, name: 'miss' }, 
  { id: 3, name: 'out' }, 
  { id: 4, name: 'your' }, 
  { id: 5, name: 'savings' }, 
  { id: 6, name: 'and' } 
];

// Helper to escape special regex characters in topics (avoids bugs with symbols like . or *)
function escapeRegExp(string) {
  return string.replace(/[.*+?^${}()|[\]\\]/g, '\\$&');
}

function fetchTopics() {
  // Step 1: Build our hash map for instant topic lookups
  const topicMap = new Map();
  // Step 2: Build a combined regex pattern from all topics
  const topicPatterns = topicData.map(datum => {
    topicMap.set(datum.name, datum.id);
    return escapeRegExp(datum.name);
  });
  // Combine into one regex with global flag to find all matches
  const combinedRegex = new RegExp(topicPatterns.join('|'), 'g');

  // Run ONE pass over the DOM with our optimized setup
  searchPage(combinedRegex, topicMap);

  // If you switch back to API fetch later, adjust here to process fetched data the same way
  // fetch(apiUrl, { method: 'GET', headers })
  //   .then(res => res.json())
  //   .then(data => {
  //     const topicMap = new Map();
  //     const patterns = data.map(datum => {
  //       topicMap.set(datum.name, datum.id);
  //       return escapeRegExp(datum.name);
  //     });
  //     searchPage(new RegExp(patterns.join('|'), 'g'), topicMap);
  //   })
  //   .catch(err => console.log('err', err));
}

function searchPage(combinedRegex, topicMap) {
  var walker = document.createTreeWalker(
    document.body,
    NodeFilter.SHOW_TEXT,
    function(node) {
      // Skip already highlighted nodes to avoid redundant work
      if (node.parentNode.classList?.contains('highlighted_text')) {
        return NodeFilter.FILTER_SKIP;
      }

      let regs;
      // Reset regex index to prevent infinite loops with global flag
      combinedRegex.lastIndex = 0;
      // Find all matches in the current text node
      while ((regs = combinedRegex.exec(node.textContent))) {
        const matchedTopic = regs[0];
        const topicId = topicMap.get(matchedTopic);

        if (topicId) {
          // Split the text node to insert the highlighted link
          var after = node.splitText(regs.index);
          after.nodeValue = after.nodeValue.substring(matchedTopic.length);
          
          var match = document.createElement('A');
          match.appendChild(document.createTextNode(matchedTopic));
          match.href = `https://app.com/CustomTopic/${topicId}`;
          match.classList.add('highlighted_text');
          
          node.parentNode.insertBefore(match, after);
          
          // Update node reference to the remaining text after the match
          node = after;
          // Reset regex index since we've modified the text node
          combinedRegex.lastIndex = 0;
        }
      }
      return NodeFilter.FILTER_SKIP;
    },
    false
  );

  // Traverse all text nodes in the DOM
  while (walker.nextNode());
}

// Initialize on page load and hash change
$(window).bind('load', function() {
  document.addEventListener('click', init); // Assuming `init` is defined elsewhere
  fetchTopics();
});

$(window).on('hashchange', function() {
  fetchTopics();
});

Key Optimizations Explained

  • Single DOM Traversal: Instead of 1000+ passes, we scan the DOM once. This eliminates the most expensive part of your original code.
  • O(1) Hash Map Lookups: The topicMap lets us fetch a topic's ID instantly when we find a match—no more looping through the entire topic array.
  • Combined Regex: Matching all topics in one regex call is far faster than running a separate regex for each topic. We also escape special characters to avoid regex bugs.
  • Avoid Reprocessing: We skip already highlighted nodes, so we don't waste time on content we've already handled.

Bonus Tips

  • If you need case-insensitive matching, add the i flag to the combined regex: new RegExp(patterns.join('|'), 'gi')
  • For extremely large topic lists (10k+), you could split the regex into smaller chunks, but 1000 topics will work perfectly with this approach.

内容的提问来源于stack exchange,提问作者milan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 07:52:10