You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Java中如何校验URL或字符串是否在域名等黑白名单内?

灵活的URL黑白名单校验方案

核心思路

直接用字符串匹配(contains()/equals())无法适配URL的多格式场景,核心解决思路是:

  1. 将输入URL和白名单规则都解析为结构化数据(拆分协议、域名、路径等维度)
  2. 根据规则的预设匹配模式(精确域名、域名前缀、完整URL、路径前缀等),针对性校验对应维度的匹配关系
  3. 避免模糊匹配带来的误判(如google.comm不会匹配google.com)

具体实现步骤

1. 定义规则匹配模式

为白名单规则预设几种常用匹配模式,可通过规则前缀约定:

  • DOMAIN_PREFIX(默认):规则为纯域名(如google.com),匹配该域名下所有路径的URL(无论协议)
  • DOMAIN_EXACT:规则前缀加=(如=google.com),仅匹配该域名的根路径URL(如http://google.com或https://google.com/,不匹配google.com/page)
  • FULL_URL_EXACT:规则带完整协议(如https://google.com),精确匹配协议+域名+根路径
  • PATH_PREFIX:规则为域名+路径(如hello.com/example),匹配该域名下以指定路径为前缀的URL(如hello.com/example/abc,不匹配hello.com/test)
  • PATH_EXACT:规则前缀加=(如=hello.com/example),精确匹配域名+完整路径

2. 解析URL为结构化对象

使用Java自带的java.net.URI类,将输入字符串解析为包含以下字段的对象:

  • 协议(protocol):如http、https
  • 域名(host):如google.com
  • 路径(path):如/example(默认根路径为/)

3. 预处理白名单规则

加载whitelist.txt时,将每条规则解析为包含匹配模式和结构化字段的规则对象:

  • 识别规则前缀(如=)确定匹配模式
  • 拆分规则的协议、域名、路径字段

4. 实现匹配逻辑

针对不同匹配模式,编写对应的校验逻辑:

  • DOMAIN_PREFIX:输入URL的域名与规则域名完全相等,忽略协议和路径
  • DOMAIN_EXACT:输入URL的域名相等,且路径为/(根路径)
  • FULL_URL_EXACT:输入URL的协议、域名、路径均与规则完全相等
  • PATH_PREFIX:域名相等,且输入路径以规则路径为前缀(需处理路径末尾的/)
  • PATH_EXACT:域名相等,且路径完全相等

代码示例

1. 定义URL结构化类

import java.net.URI;
import java.net.URISyntaxException;

public class ParsedUrl {
    private String protocol;
    private String host;
    private String path;

    public ParsedUrl(String urlStr) throws URISyntaxException {
        URI uri = new URI(urlStr.startsWith("http") ? urlStr : "http://" + urlStr);
        this.protocol = uri.getScheme() != null ? uri.getScheme() : "http";
        this.host = uri.getHost() != null ? uri.getHost() : uri.getPath().split("/")[0];
        this.path = uri.getPath() != null ? uri.getPath() : "/";
        // 处理空路径为根路径
        if (this.path.isEmpty()) {
            this.path = "/";
        }
    }

    // Getters
    public String getProtocol() { return protocol; }
    public String getHost() { return host; }
    public String getPath() { return path; }
}

2. 定义白名单规则类

public class WhitelistRule {
    public enum MatchMode {
        DOMAIN_PREFIX, DOMAIN_EXACT, FULL_URL_EXACT, PATH_PREFIX, PATH_EXACT
    }

    private MatchMode mode;
    private String protocol;
    private String host;
    private String path;

    public WhitelistRule(String ruleStr) throws URISyntaxException {
        boolean isExact = ruleStr.startsWith("=");
        String cleanRule = isExact ? ruleStr.substring(1) : ruleStr;

        // 识别匹配模式
        if (cleanRule.startsWith("http://") || cleanRule.startsWith("https://")) {
            this.mode = isExact ? MatchMode.FULL_URL_EXACT : MatchMode.FULL_URL_EXACT;
            ParsedUrl parsedRule = new ParsedUrl(cleanRule);
            this.protocol = parsedRule.getProtocol();
            this.host = parsedRule.getHost();
            this.path = parsedRule.getPath();
        } else if (cleanRule.contains("/")) {
            this.mode = isExact ? MatchMode.PATH_EXACT : MatchMode.PATH_PREFIX;
            // 拆分域名和路径
            String[] parts = cleanRule.split("/", 2);
            this.host = parts[0];
            this.path = "/" + parts[1];
            // 统一路径末尾的/
            if (!this.path.endsWith("/")) {
                this.path += "/";
            }
        } else {
            this.mode = isExact ? MatchMode.DOMAIN_EXACT : MatchMode.DOMAIN_PREFIX;
            this.host = cleanRule;
            this.path = "/";
        }
    }

    // 匹配逻辑
    public boolean matches(ParsedUrl inputUrl) {
        switch (mode) {
            case DOMAIN_PREFIX:
                return inputUrl.getHost().equals(this.host);
            case DOMAIN_EXACT:
                return inputUrl.getHost().equals(this.host) && inputUrl.getPath().equals("/");
            case FULL_URL_EXACT:
                return inputUrl.getProtocol().equals(this.protocol)
                        && inputUrl.getHost().equals(this.host)
                        && inputUrl.getPath().equals(this.path);
            case PATH_PREFIX:
                if (!inputUrl.getHost().equals(this.host)) return false;
                String inputPath = inputUrl.getPath().endsWith("/") ? inputUrl.getPath() : inputUrl.getPath() + "/";
                return inputPath.startsWith(this.path);
            case PATH_EXACT:
                return inputUrl.getHost().equals(this.host) && inputUrl.getPath().equals(this.path);
            default:
                return false;
        }
    }
}

3. 加载白名单并校验

import java.io.IOException;
import java.nio.file.Files;
import java.nio.file.Paths;
import java.util.HashSet;
import java.util.Set;

public class WhitelistValidator {
    private Set<WhitelistRule> rules = new HashSet<>();

    public WhitelistValidator(String whitelistPath) throws IOException, URISyntaxException {
        // 加载白名单文件
        for (String line : Files.readAllLines(Paths.get(whitelistPath))) {
            line = line.trim();
            if (!line.isEmpty()) {
                rules.add(new WhitelistRule(line));
            }
        }
    }

    public boolean isWhitelisted(String urlStr) {
        try {
            ParsedUrl inputUrl = new ParsedUrl(urlStr);
            for (WhitelistRule rule : rules) {
                if (rule.matches(inputUrl)) {
                    return true;
                }
            }
        } catch (URISyntaxException e) {
            // 非法URL视为不匹配
            return false;
        }
        return false;
    }

    // 测试示例
    public static void main(String[] args) throws IOException, URISyntaxException {
        WhitelistValidator validator = new WhitelistValidator("whitelist.txt");
        // 测试用例
        System.out.println(validator.isWhitelisted("http://google.com")); // true
        System.out.println(validator.isWhitelisted("https://google.com")); // true
        System.out.println(validator.isWhitelisted("google.comm")); // false
        System.out.println(validator.isWhitelisted("hello.com")); // false
        System.out.println(validator.isWhitelisted("hello.com/example")); // true
        System.out.println(validator.isWhitelisted("hello.com/example/test")); // true
    }
}

扩展适配

  • 黑名单适配:只需将isWhitelisted方法的逻辑反转(匹配任意黑名单规则则返回false)
  • 端口支持:可在ParsedUrl中加入端口字段,匹配时根据规则是否包含端口决定是否校验
  • 通配符支持:若需*.google.com这类规则,可引入正则表达式匹配域名部分

内容的提问来源于stack exchange,提问作者James Bayhaner

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 22:27:06