如何编写正则表达式:识别非a-z0-9字符及Django正则路由
Alright, let's break down your two questions one by one—happy to help with these regex and Django routing challenges!
1. Regex to Detect Characters Outside a-z/0-9
To spot if a string contains any character that’s not a lowercase letter (a-z) or digit (0-9), use a negated character class—this is the most straightforward approach:
[^a-z0-9]
- The
^inside square brackets negates the class, so it matches any character not in thea-z0-9set. - If you want to verify the entire string contains at least one such character (rather than just matching the character itself), wrap the class in wildcards to cover surrounding text:
.*[^a-z0-9].* - Add the
imodifier (case-insensitive) if you want to allow uppercase letters (A-Z) too:[^a-zA-Z0-9]
Quick Python Example:
import re def has_non_alnum_char(s): return bool(re.search(r'[^a-z0-9]', s)) # Test cases print(has_non_alnum_char("hello123")) # False (only a-z/0-9) print(has_non_alnum_char("hello!123")) # True (contains "!")
2. Detect Regex Syntax in Django URL Routes
For your Django middleware use case—identifying if a request’s associated route uses regex (instead of plain string paths)—you need to target the regex-specific syntax common in re_path() (or legacy url()) definitions.
Key Regex Elements to Target
Django regex routes almost always include one or more of these:
- Standard regex meta-characters:
^,$,.,*,+,?,(),[],{},|,\ - Django-specific named capture groups:
(?P<name>...) - Numeric quantifiers like
{4}(e.g., for matching 4-digit years)
The Detection Regex
Use this pattern to match any of these regex-specific elements:
[\\^$.*+?()[\]{}|]|(?P<\w+>)|\{\d+\}
[\\^$.*+?()[\]{}|]: Matches escaped regex meta-characters (we escape\because it’s a special character in regex strings)(?P<\w+>): Catches Django’s named capture group syntax\{\d+\}: Identifies numeric quantifiers like{4}
Django Middleware Implementation Example
import re from django.http import HttpResponse from django.urls import resolve class RegexURLInterceptorMiddleware: def __init__(self, get_response): self.get_response = get_response def __call__(self, request): # Resolve the request path to its associated route pattern try: match = resolve(request.path) route_pattern = match.pattern.regex.pattern except: # Handle paths that don't match any route route_pattern = "" # Check if the route uses regex syntax has_regex_syntax = bool(re.search(r'[\\^$.*+?()[\]{}|]|(?P<\w+>)', route_pattern)) if has_regex_syntax: # Your logic to parse urls.py and standardize the URL goes here print(f"Intercepted regex route: {route_pattern}") # Example: Modify the request path to a standardized format # request.path = self.standardize_url(request.path, route_pattern) response = self.get_response(request) return response def standardize_url(self, path, route_pattern): # Implement your URL standardization logic here # For example, replace regex captures with standardized placeholders return path
Important Notes
- Django’s
path()uses path converters (like<int:pk>) which aren’t raw regex, but they translate to regex under the hood. If you want to exclude these from your "regex route" check, adjust the pattern to ignore<\w+:\w+>syntax. - Using Django’s built-in
resolve()function is far more reliable than parsing urls.py directly—it gives you the exact regex pattern associated with the request path.
内容的提问来源于stack exchange,提问作者Grant Zukel
相关产品推荐
相关产品推荐

