Java中如何校验URL或字符串是否在域名等黑白名单内?
灵活的URL黑白名单校验方案
核心思路
直接用字符串匹配(contains()/equals())无法适配URL的多格式场景,核心解决思路是:
- 将输入URL和白名单规则都解析为结构化数据(拆分协议、域名、路径等维度)
- 根据规则的预设匹配模式(精确域名、域名前缀、完整URL、路径前缀等),针对性校验对应维度的匹配关系
- 避免模糊匹配带来的误判(如
google.comm不会匹配google.com)
具体实现步骤
1. 定义规则匹配模式
为白名单规则预设几种常用匹配模式,可通过规则前缀约定:
DOMAIN_PREFIX(默认):规则为纯域名(如google.com),匹配该域名下所有路径的URL(无论协议)DOMAIN_EXACT:规则前缀加=(如=google.com),仅匹配该域名的根路径URL(如http://google.com或https://google.com/,不匹配google.com/page)FULL_URL_EXACT:规则带完整协议(如https://google.com),精确匹配协议+域名+根路径PATH_PREFIX:规则为域名+路径(如hello.com/example),匹配该域名下以指定路径为前缀的URL(如hello.com/example/abc,不匹配hello.com/test)PATH_EXACT:规则前缀加=(如=hello.com/example),精确匹配域名+完整路径
2. 解析URL为结构化对象
使用Java自带的java.net.URI类,将输入字符串解析为包含以下字段的对象:
- 协议(protocol):如
http、https - 域名(host):如
google.com - 路径(path):如
/example(默认根路径为/)
3. 预处理白名单规则
加载whitelist.txt时,将每条规则解析为包含匹配模式和结构化字段的规则对象:
- 识别规则前缀(如
=)确定匹配模式 - 拆分规则的协议、域名、路径字段
4. 实现匹配逻辑
针对不同匹配模式,编写对应的校验逻辑:
- DOMAIN_PREFIX:输入URL的域名与规则域名完全相等,忽略协议和路径
- DOMAIN_EXACT:输入URL的域名相等,且路径为
/(根路径) - FULL_URL_EXACT:输入URL的协议、域名、路径均与规则完全相等
- PATH_PREFIX:域名相等,且输入路径以规则路径为前缀(需处理路径末尾的
/) - PATH_EXACT:域名相等,且路径完全相等
代码示例
1. 定义URL结构化类
import java.net.URI; import java.net.URISyntaxException; public class ParsedUrl { private String protocol; private String host; private String path; public ParsedUrl(String urlStr) throws URISyntaxException { URI uri = new URI(urlStr.startsWith("http") ? urlStr : "http://" + urlStr); this.protocol = uri.getScheme() != null ? uri.getScheme() : "http"; this.host = uri.getHost() != null ? uri.getHost() : uri.getPath().split("/")[0]; this.path = uri.getPath() != null ? uri.getPath() : "/"; // 处理空路径为根路径 if (this.path.isEmpty()) { this.path = "/"; } } // Getters public String getProtocol() { return protocol; } public String getHost() { return host; } public String getPath() { return path; } }
2. 定义白名单规则类
public class WhitelistRule { public enum MatchMode { DOMAIN_PREFIX, DOMAIN_EXACT, FULL_URL_EXACT, PATH_PREFIX, PATH_EXACT } private MatchMode mode; private String protocol; private String host; private String path; public WhitelistRule(String ruleStr) throws URISyntaxException { boolean isExact = ruleStr.startsWith("="); String cleanRule = isExact ? ruleStr.substring(1) : ruleStr; // 识别匹配模式 if (cleanRule.startsWith("http://") || cleanRule.startsWith("https://")) { this.mode = isExact ? MatchMode.FULL_URL_EXACT : MatchMode.FULL_URL_EXACT; ParsedUrl parsedRule = new ParsedUrl(cleanRule); this.protocol = parsedRule.getProtocol(); this.host = parsedRule.getHost(); this.path = parsedRule.getPath(); } else if (cleanRule.contains("/")) { this.mode = isExact ? MatchMode.PATH_EXACT : MatchMode.PATH_PREFIX; // 拆分域名和路径 String[] parts = cleanRule.split("/", 2); this.host = parts[0]; this.path = "/" + parts[1]; // 统一路径末尾的/ if (!this.path.endsWith("/")) { this.path += "/"; } } else { this.mode = isExact ? MatchMode.DOMAIN_EXACT : MatchMode.DOMAIN_PREFIX; this.host = cleanRule; this.path = "/"; } } // 匹配逻辑 public boolean matches(ParsedUrl inputUrl) { switch (mode) { case DOMAIN_PREFIX: return inputUrl.getHost().equals(this.host); case DOMAIN_EXACT: return inputUrl.getHost().equals(this.host) && inputUrl.getPath().equals("/"); case FULL_URL_EXACT: return inputUrl.getProtocol().equals(this.protocol) && inputUrl.getHost().equals(this.host) && inputUrl.getPath().equals(this.path); case PATH_PREFIX: if (!inputUrl.getHost().equals(this.host)) return false; String inputPath = inputUrl.getPath().endsWith("/") ? inputUrl.getPath() : inputUrl.getPath() + "/"; return inputPath.startsWith(this.path); case PATH_EXACT: return inputUrl.getHost().equals(this.host) && inputUrl.getPath().equals(this.path); default: return false; } } }
3. 加载白名单并校验
import java.io.IOException; import java.nio.file.Files; import java.nio.file.Paths; import java.util.HashSet; import java.util.Set; public class WhitelistValidator { private Set<WhitelistRule> rules = new HashSet<>(); public WhitelistValidator(String whitelistPath) throws IOException, URISyntaxException { // 加载白名单文件 for (String line : Files.readAllLines(Paths.get(whitelistPath))) { line = line.trim(); if (!line.isEmpty()) { rules.add(new WhitelistRule(line)); } } } public boolean isWhitelisted(String urlStr) { try { ParsedUrl inputUrl = new ParsedUrl(urlStr); for (WhitelistRule rule : rules) { if (rule.matches(inputUrl)) { return true; } } } catch (URISyntaxException e) { // 非法URL视为不匹配 return false; } return false; } // 测试示例 public static void main(String[] args) throws IOException, URISyntaxException { WhitelistValidator validator = new WhitelistValidator("whitelist.txt"); // 测试用例 System.out.println(validator.isWhitelisted("http://google.com")); // true System.out.println(validator.isWhitelisted("https://google.com")); // true System.out.println(validator.isWhitelisted("google.comm")); // false System.out.println(validator.isWhitelisted("hello.com")); // false System.out.println(validator.isWhitelisted("hello.com/example")); // true System.out.println(validator.isWhitelisted("hello.com/example/test")); // true } }
扩展适配
- 黑名单适配:只需将
isWhitelisted方法的逻辑反转(匹配任意黑名单规则则返回false) - 端口支持:可在
ParsedUrl中加入端口字段,匹配时根据规则是否包含端口决定是否校验 - 通配符支持:若需
*.google.com这类规则,可引入正则表达式匹配域名部分
内容的提问来源于stack exchange,提问作者James Bayhaner
相关产品推荐
相关产品推荐

