JavaScript子域名验证正则表达式需求:排除顶级域名
Got it, let's break down what you need: you want a regex that rejects "base" registered domains like domain.uk or domain.ac.uk, but accepts their subdomains like sub.domain.uk or sub.domain.ac.uk.
First, let's clarify the core rule from your examples:
- Reject domains that are directly registered under a public suffix (like
.ukor.ac.uk) - Accept domains that have an extra subdomain prefix added to those registered base domains
The challenge here is that public suffixes can be single-segment (.uk, .com) or multi-segment (.ac.uk, .com.au). A pure regex can't perfectly cover every possible public suffix (since the list is maintained by organizations like Mozilla), but we can build a solid approximation that works for your examples and most common cases.
The Regex
Here's a tailored regex for JavaScript:
const subdomainRegex = /^(?!^[a-zA-Z0-9-]+\.[a-zA-Z0-9-]{2,}(\.[a-zA-Z0-9-]{2,})?$)[a-zA-Z0-9-]+(\.[a-zA-Z0-9-]+){2,}$/;
How It Works
Let's break down the components:
Negative Lookahead (
(?!...)): This part rejects any domain that matches the "base registered domain" pattern:^[a-zA-Z0-9-]+: The main domain label (e.g.,domainindomain.uk)\.[a-zA-Z0-9-]{2,}: A top-level suffix (e.g.,.uk,.com)(\.[a-zA-Z0-9-]{2,})?$: Optional second segment for multi-part suffixes (e.g.,.acin.ac.uk)
This ensures we exclude domains likedomain.uk(2 segments) anddomain.ac.uk(3 segments).
Main Matching Pattern:
[a-zA-Z0-9-]+: The subdomain label (e.g.,subinsub.domain.uk)(\.[a-zA-Z0-9-]+){2,}: At least two additional segments (the base registered domain, which can be 2 or more segments)
This ensures we accept domains likesub.domain.uk(3 segments) andsub.domain.ac.uk(4 segments).
Testing It Out
Let's verify with your examples:
domain.uk→ ❌ Rejected (matches the base domain pattern)sub.domain.uk→ ✅ Accepteddomain.ac.uk→ ❌ Rejected (matches the base domain pattern)sub.domain.ac.uk→ ✅ Accepted
Notes on Limitations
- This regex uses a general pattern for public suffixes. If you need to cover rare or custom multi-segment suffixes, you'd need to expand the negative lookahead to target those specific cases.
- It doesn't handle edge cases like internationalized domain names (IDNs) unless you add support for Unicode characters (using
\p{L}instead ofa-zA-Z, but note that JavaScript regex requires theuflag for Unicode property escapes).
Example Usage
function isSubdomain(hostname) { const subdomainRegex = /^(?!^[a-zA-Z0-9-]+\.[a-zA-Z0-9-]{2,}(\.[a-zA-Z0-9-]{2,})?$)[a-zA-Z0-9-]+(\.[a-zA-Z0-9-]+){2,}$/; return subdomainRegex.test(hostname); } // Test cases console.log(isSubdomain("domain.uk")); // false console.log(isSubdomain("sub.domain.uk")); // true console.log(isSubdomain("domain.ac.uk")); // false console.log(isSubdomain("sub.domain.ac.uk")); // true
内容的提问来源于stack exchange,提问作者HA Luca

