如何用grep/sed精准提取顶级域名?grep新手技术求助
Hey there! I totally get the frustration of running a grep command only to have to manually trim the results afterward—especially when you're still getting comfortable with regex. Let's fix that so you can pull the exact domain you need in one go.
Looking at your input examples (like //cdn.www2.qqw-ab-mon.com), it seems you want the main domain plus top-level domain (e.g., qqw-ab-mon.com) instead of the full subdomain path. Here are a few clean approaches:
Method 1: Grep + Grep (Double Filter)
First, we'll extract everything after // up to the closing quote, then narrow it down to the final domain segment:
grep -Eo '//[^"]+' file | grep -Eo '[a-zA-Z0-9-]+\.[a-zA-Z]{2,}$'
- The first
grepgrabs all text starting with//until it hits a"(since your input has domains wrapped in quotes likewebsiteHost:"//..."). - The second
greptargets the last part of that string: it matches a segment of letters/numbers/dashes, followed by a dot, then a 2+ letter top-level domain (like.com).
Method 2: Grep + Awk (Simpler for Standard TLDs)
If your domains all use standard 2-letter TLDs (like .com, .org), awk makes this super clean by splitting the string on dots and grabbing the last two segments:
grep -Eo '//[^"]+' file | awk -F. '{print $(NF-1)"."$NF}'
-F.tells awk to split each line by dots.$(NF-1)is the second-to-last field, and$NFis the last field—combining them gives you the main domain + TLD.
Bonus: One-Liner Grep (For Regex Pros)
If you want to do it all in a single grep command, you can use a lookbehind to target the right part of the string:
grep -Eo '(?<=//)[^"]*' file | grep -Eo '[a-zA-Z0-9-]+\.[a-zA-Z]{2,}$'
The (?<=//) is a positive lookbehind—it ensures we only match text that comes right after //, without including the // itself.
Example Output
For your input lines like:
"asynch",websiteHost:"//cdn.www2.qqw-ab-mon.com",Myhost:"// "
All these methods will output:
qqw-ab-mon.com
No more extra trimming—done in one or two quick commands!
内容的提问来源于stack exchange,提问作者Bridget Roe

