如何利用哈希表优化Gmail中千级主题文本匹配的性能
Fixing Gmail Performance Issues with 1000+ Topic Matches
Let's break down why your current code is causing Gmail to choke, then fix it with a hash map (and smart regex optimization) to get things running smoothly.
The Root of the Problem
Your current approach runs searchPage() once per topic—that's 1000+ separate passes over Gmail's massive DOM tree. Each pass uses a TreeWalker and regex, which adds up to way too much work for the browser. No wonder it's crashing!
The Fix: One DOM Pass, All Matches
Instead of looping through each topic and re-scanning the DOM every time, we'll:
- Combine all topics into a single regex (so we can match every topic in one go)
- Use a
Map(our hash table) to quickly look up topic data when we find a match - Only traverse the DOM once to find and highlight all matches
Here's the optimized code:
var topicData = [ { id: 1, name: 'Courses' }, { id: 2, name: 'miss' }, { id: 3, name: 'out' }, { id: 4, name: 'your' }, { id: 5, name: 'savings' }, { id: 6, name: 'and' } ]; // Helper to escape special regex characters in topics (avoids bugs with symbols like . or *) function escapeRegExp(string) { return string.replace(/[.*+?^${}()|[\]\\]/g, '\\$&'); } function fetchTopics() { // Step 1: Build our hash map for instant topic lookups const topicMap = new Map(); // Step 2: Build a combined regex pattern from all topics const topicPatterns = topicData.map(datum => { topicMap.set(datum.name, datum.id); return escapeRegExp(datum.name); }); // Combine into one regex with global flag to find all matches const combinedRegex = new RegExp(topicPatterns.join('|'), 'g'); // Run ONE pass over the DOM with our optimized setup searchPage(combinedRegex, topicMap); // If you switch back to API fetch later, adjust here to process fetched data the same way // fetch(apiUrl, { method: 'GET', headers }) // .then(res => res.json()) // .then(data => { // const topicMap = new Map(); // const patterns = data.map(datum => { // topicMap.set(datum.name, datum.id); // return escapeRegExp(datum.name); // }); // searchPage(new RegExp(patterns.join('|'), 'g'), topicMap); // }) // .catch(err => console.log('err', err)); } function searchPage(combinedRegex, topicMap) { var walker = document.createTreeWalker( document.body, NodeFilter.SHOW_TEXT, function(node) { // Skip already highlighted nodes to avoid redundant work if (node.parentNode.classList?.contains('highlighted_text')) { return NodeFilter.FILTER_SKIP; } let regs; // Reset regex index to prevent infinite loops with global flag combinedRegex.lastIndex = 0; // Find all matches in the current text node while ((regs = combinedRegex.exec(node.textContent))) { const matchedTopic = regs[0]; const topicId = topicMap.get(matchedTopic); if (topicId) { // Split the text node to insert the highlighted link var after = node.splitText(regs.index); after.nodeValue = after.nodeValue.substring(matchedTopic.length); var match = document.createElement('A'); match.appendChild(document.createTextNode(matchedTopic)); match.href = `https://app.com/CustomTopic/${topicId}`; match.classList.add('highlighted_text'); node.parentNode.insertBefore(match, after); // Update node reference to the remaining text after the match node = after; // Reset regex index since we've modified the text node combinedRegex.lastIndex = 0; } } return NodeFilter.FILTER_SKIP; }, false ); // Traverse all text nodes in the DOM while (walker.nextNode()); } // Initialize on page load and hash change $(window).bind('load', function() { document.addEventListener('click', init); // Assuming `init` is defined elsewhere fetchTopics(); }); $(window).on('hashchange', function() { fetchTopics(); });
Key Optimizations Explained
- Single DOM Traversal: Instead of 1000+ passes, we scan the DOM once. This eliminates the most expensive part of your original code.
- O(1) Hash Map Lookups: The
topicMaplets us fetch a topic's ID instantly when we find a match—no more looping through the entire topic array. - Combined Regex: Matching all topics in one regex call is far faster than running a separate regex for each topic. We also escape special characters to avoid regex bugs.
- Avoid Reprocessing: We skip already highlighted nodes, so we don't waste time on content we've already handled.
Bonus Tips
- If you need case-insensitive matching, add the
iflag to the combined regex:new RegExp(patterns.join('|'), 'gi') - For extremely large topic lists (10k+), you could split the regex into smaller chunks, but 1000 topics will work perfectly with this approach.
内容的提问来源于stack exchange,提问作者milan
相关产品推荐
相关产品推荐

