Python正则搜索函数性能优化:单请求调用超45万次
search Function Hey there! With your search function getting called over 452,900 times per request, even tiny optimizations will have a huge impact on overall performance. Let's walk through the bottlenecks in your current code and fix them step by step.
First, Let's Break Down the Original Code's Pain Points
def search(self, arg1, arg2):
global ipv4_regex, type_regex
if isinstance(arg2, type_regex):
text = re.search(arg2, arg1)
if text:
p = text.group()
return len(p)
else:
return 0
else:
if arg2 in arg1:
return len(arg2)
else:
return 0
The main issues holding back performance here are:
- Unnecessary global variable access (you only use
type_regex, and don't modify it, so theglobaldeclaration is redundant and slower than local access) - Suboptimal branch order (if string matches are more common, we're wasting cycles checking regex first)
- Redundant string creation when calculating regex match length
- Slightly slower
incheck for string presence
Optimized Code
Here's the revised version with key performance tweaks:
def search(self, arg1, arg2): # Prioritize string matching (swap order if regex matches are more frequent) if not isinstance(arg2, type_regex): # Use str.find() instead of 'in' for faster presence checks if arg1.find(arg2) != -1: return len(arg2) return 0 # Handle regex case with minimal overhead match = re.search(arg2, arg1) # Calculate length directly without extracting the full match string return match.end() - match.start() if match else 0
Breakdown of Each Optimization
- Reordered branches: If most of your calls are string-based matches (the
elsecase in your original code), we skip the regex type check entirely for those calls. This cuts down on unnecessaryisinstancechecks across 450k+ invocations. - Replaced
inwithstr.find():str.find()returns an index (or-1if not found) and has slightly lower overhead than theinoperator in high-frequency scenarios. - Avoided string extraction: Instead of pulling the full matched string with
.group()and then getting its length, we calculate the length directly usingmatch.end() - match.start(). This skips creating a new string object entirely—big win for both memory and speed. - Removed redundant
globaldeclaration: Since you're only readingtype_regex(not modifying it), thegloballine was unnecessary. Global variable access is slower than local access, so removing this saves a tiny bit of time per call. - Simplified return logic: We consolidated multiple
returnstatements to reduce interpreter jumps, making the code more compact and efficient.
Advanced Optimizations to Squeeze Even More Performance
If you want to push this further:
- Cache repeated results: Use
functools.lru_cacheto cache results for repeatedarg1/arg2pairs. Just make sure your arguments are hashable (strings and precompiled regexes are):from functools import lru_cache # Note: If your `self` instance affects the function's behavior, adjust this (e.g., make it static) @lru_cache(maxsize=2048) # Adjust maxsize based on your memory constraints def search(self, arg1, arg2): # Optimized code from above... - Confirm regexes are precompiled: Ensure all
type_regexinstances are created withre.compile()—dynamic regex compilation is a massive performance drain, and your current code assumes these are precompiled, but it's worth double-checking.
Once you've implemented these changes, run your pprofile again to measure the speedup—you should see a noticeable drop in the time spent in this function.
内容的提问来源于stack exchange,提问作者prashantgpt91

