You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python正则搜索函数性能优化:单请求调用超45万次

Optimizing Your High-Call-Count search Function

Hey there! With your search function getting called over 452,900 times per request, even tiny optimizations will have a huge impact on overall performance. Let's walk through the bottlenecks in your current code and fix them step by step.

First, Let's Break Down the Original Code's Pain Points

def search(self, arg1, arg2):
global ipv4_regex, type_regex
if isinstance(arg2, type_regex):
text = re.search(arg2, arg1)
if text:
p = text.group()
return len(p)
else:
return 0
else:
if arg2 in arg1:
return len(arg2)
else:
return 0

The main issues holding back performance here are:

  • Unnecessary global variable access (you only use type_regex, and don't modify it, so the global declaration is redundant and slower than local access)
  • Suboptimal branch order (if string matches are more common, we're wasting cycles checking regex first)
  • Redundant string creation when calculating regex match length
  • Slightly slower in check for string presence

Optimized Code

Here's the revised version with key performance tweaks:

def search(self, arg1, arg2):
    # Prioritize string matching (swap order if regex matches are more frequent)
    if not isinstance(arg2, type_regex):
        # Use str.find() instead of 'in' for faster presence checks
        if arg1.find(arg2) != -1:
            return len(arg2)
        return 0
    
    # Handle regex case with minimal overhead
    match = re.search(arg2, arg1)
    # Calculate length directly without extracting the full match string
    return match.end() - match.start() if match else 0

Breakdown of Each Optimization

  • Reordered branches: If most of your calls are string-based matches (the else case in your original code), we skip the regex type check entirely for those calls. This cuts down on unnecessary isinstance checks across 450k+ invocations.
  • Replaced in with str.find(): str.find() returns an index (or -1 if not found) and has slightly lower overhead than the in operator in high-frequency scenarios.
  • Avoided string extraction: Instead of pulling the full matched string with .group() and then getting its length, we calculate the length directly using match.end() - match.start(). This skips creating a new string object entirely—big win for both memory and speed.
  • Removed redundant global declaration: Since you're only reading type_regex (not modifying it), the global line was unnecessary. Global variable access is slower than local access, so removing this saves a tiny bit of time per call.
  • Simplified return logic: We consolidated multiple return statements to reduce interpreter jumps, making the code more compact and efficient.

Advanced Optimizations to Squeeze Even More Performance

If you want to push this further:

  • Cache repeated results: Use functools.lru_cache to cache results for repeated arg1/arg2 pairs. Just make sure your arguments are hashable (strings and precompiled regexes are):
    from functools import lru_cache
    
    # Note: If your `self` instance affects the function's behavior, adjust this (e.g., make it static)
    @lru_cache(maxsize=2048)  # Adjust maxsize based on your memory constraints
    def search(self, arg1, arg2):
        # Optimized code from above...
    
  • Confirm regexes are precompiled: Ensure all type_regex instances are created with re.compile()—dynamic regex compilation is a massive performance drain, and your current code assumes these are precompiled, but it's worth double-checking.

Once you've implemented these changes, run your pprofile again to measure the speedup—you should see a noticeable drop in the time spent in this function.

内容的提问来源于stack exchange,提问作者prashantgpt91

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:55:20