求助:使用正则表达式剥离URL前缀提取核心域名
Solution to Extract Core Domains from WWW-Prefixed URLs
Hey there! Let's tackle this problem. The key here is to create a regex pattern that can handle all the variations of your www-like prefixes—whether they have numbers, hyphens, or are directly attached to the domain without a dot.
Here's a Python solution that works for all your examples:
import re examples = ['www2.chccs.k12.nc.us', 'wwwsco.com', 'www-152.aig.com', 'www.google.com'] core_domains = [re.sub(r'^www(?:[\d-]+)?\.?', '', domain) for domain in examples] print(core_domains) # Output: ['chccs.k12.nc.us', 'sco.com', 'aig.com', 'google.com']
Breaking Down the Regex: ^www(?:[\d-]+)?\.?
Let's break down what each part does to understand why it works:
^: Anchors the match to the start of the string (so we only target leading prefixes).www: Matches the literal "www" at the beginning.(?:[\d-]+)?: An optional non-capturing group that matches one or more digits (\d) or hyphens (-). The?makes this group optional, so it handles cases where there's nothing after "www" (like yourwwwsco.comexample).\.?: Matches an optional dot. This ensures we remove the dot along with the prefix when it exists (likewww.google.com→ removewww.instead of justwww).
Edge Cases to Consider
- If you have a domain like
www-abc.com, this regex will removewww-leavingabc.com—which aligns with your example logic for hyphenated prefixes. - For domains without any www prefix, the regex will leave them untouched (since it only matches leading www patterns).
This approach is clean and handles all the scenarios you provided perfectly.
内容的提问来源于stack exchange,提问作者Reid
相关产品推荐
相关产品推荐

