You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

C# 如何从各类URL中提取包含通配符的合规有效域名

C# 带通配符域名提取实现方案

实现思路

  1. 先清洗输入内容:移除 http/https 等协议头,再截断第一个/、?、#之后的路径、参数、锚点内容,得到域名候选串
  2. 合法性校验:严格匹配通配符规则+标准域名规范:
    • 通配符*最多只能出现1次,且仅允许出现在字符串开头
    • 除通配符外,仅允许出现字母、数字、半角连接符-、半角点.
    • 域名各级段首尾不能是-,不允许出现连续的.或-
    • 顶级域最少2个字符,只能是字母

完整实现代码

using System;
using System.Text.RegularExpressions;

public static class DomainExtractor
{
    // 预编译正则,提升重复调用性能
    private static readonly Regex ValidDomainRegex = new Regex(
        @"^(\*)?([a-zA-Z0-9]([a-zA-Z0-9\-]{0,61}[a-zA-Z0-9])?\.)+[a-zA-Z]{2,}$", 
        RegexOptions.Compiled);

    public static string ExtractDomainFromUrl(string input)
    {
        if (string.IsNullOrWhiteSpace(input))
            return null;
        
        var cleanedInput = input.Trim().ToLowerInvariant();
        
        // 移除协议头
        if (cleanedInput.Contains("://"))
        {
            var protocolSplit = cleanedInput.Split(new[] {"://"}, StringSplitOptions.RemoveEmptyEntries);
            if (protocolSplit.Length < 2) 
                return null;
            cleanedInput = protocolSplit[1];
        }
        
        // 移除路径、查询参数、锚点
        var separatorIndex = cleanedInput.IndexOfAny(new[] {'/', '?', '#'});
        if (separatorIndex >= 0)
        {
            cleanedInput = cleanedInput.Substring(0, separatorIndex);
        }
        
        // 校验域名合法性
        return ValidDomainRegex.IsMatch(cleanedInput) ? cleanedInput : null;
    }
}

适配说明

  • 已覆盖所有给定的可接受提取场景,所有不可接受场景都会返回null
  • 支持不带协议头的输入解析,支持开头带单个通配符的域名提取
  • 如果需要适配数字顶级域(如.xyz、.123),可以把正则末尾的[a-zA-Z]{2,}修改为[a-zA-Z0-9]{2,}

内容的提问来源于stack exchange,提问作者prasana kannan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.02 18:48:02