Java文本URL检测异常:特殊格式URL无法完整识别求助
问题描述
使用Java结合com.linkedin.urls.detection.UrlDetector与org.apache.commons.validator.routines.UrlValidator实现文本URL检测功能时,可正常识别bit.ly/653282、http://google.com这类常规URL,但对于http://google/element8441876860/promethium/7893214560.com这种特殊格式的URL,仅能提取出7893214560.com,无法完整识别整个URL,需解决该问题。
相关代码
import com.linkedin.urls.detection.UrlDetector; import org.apache.commons.validator.routines.UrlValidator; boolean isValidURL(String url) { if (urlValidator == null) { System.out.println("Creating new URL Validator"); urlValidator = new UrlValidator(UrlValidator.ALLOW_2_SLASHES + UrlValidator.ALLOW_ALL_SCHEMES); } return urlValidator.isValid(url) || urlValidator.isValid("http://" + url); } // Extract and return valid URL from whole SMS text. private List<String> getURLFromSMSText(String smsTextMessage) { // smsTextMessage = "PARADISE HALEEM @Rs.249 & 1 KG FAMILY PACK @Rs.699. Your Favourite " + // "1 KG PARADISE Nizami Chicken Biryani @Rs.279 " + // "in Hyderabad Valid in Take Away " + // "- http://google/element8441876860/promethium/7893214560.com " + " bit.ly/653282 " + // "sdfsdfsf http://google/googlet.com 7893214560.com" String validUrl = null; String tempUrl = null; List<String> urlsInSms = new ArrayList<>(); try { UrlDetector parser = new UrlDetector(smsTextMessage, UrlDetectorOptions.Default); List<Url> found = parser.detect(); if (found != null && found.size() > 0) { for (Url url : found) { tempUrl = url.getOriginalUrl(); if (tempUrl != null && isValidURL(tempUrl)) { System.out.println("Valid Original URL: " + url.getOriginalUrl()); System.out.println("Scheme: " + url.getScheme()); System.out.println("Host: " + url.getHost()); System.out.println("Path: " + url.getPath()); urlsInSms.add(tempUrl); /*validUrl = tempUrl; break;*/ } else { System.out.println("Retrived Url is In-Valid (" + url.getOriginalUrl() + ")"); if (DebugConfig.getInstance().isDebugEnable()) { log.error("Retrived Url is In-Valid (" + url.getOriginalUrl() + ")"); } } } } else { if (DebugConfig.getInstance().isDebugEnable()) log.error("No Matching URL present in SMS (" + smsTextMessage + ")"); } } catch (Exception ue) { if (DebugConfig.getInstance().isDebugEnable()) log.error("Exception in Url Detection"); } return urlsInSms; }
问题原因分析
- UrlDetector拆分URL:默认的
UrlDetectorOptions.Default规则会将路径中的7893214560.com识别为独立URL,因为它符合域名格式,从而拆分了原本完整的长URL。 - UrlValidator验证不通过:
http://google/...的主机为google,无顶级域名,默认的UrlValidator会判定该URL无效,导致即使UrlDetector检测到完整URL,也会被过滤,最终只留下独立的7893214560.com。
解决方案
1. 调整UrlDetector检测规则,避免拆分完整URL
自定义UrlDetectorOptions,限制仅识别带Scheme的URL,避免将路径中的域名格式字符串当作独立URL:
UrlDetectorOptions options = UrlDetectorOptions.builder() .allowInsecureUrls(true) .allowProtocolRelativeUrls(false) // 禁用相对协议URL的检测 .allowHostlessUrls(false) // 禁用无主机URL的检测 .build(); UrlDetector parser = new UrlDetector(smsTextMessage, options);
2. 修改UrlValidator配置,允许非标准主机名
自定义主机验证逻辑,允许无顶级域名的主机通过验证:
// 自定义主机验证器 HostNameValidator customHostValidator = host -> host != null && !host.isEmpty(); // 初始化UrlValidator时使用自定义验证器 urlValidator = new UrlValidator( new String[]{"http", "https"}, // 指定允许的协议 UrlValidator.ALLOW_2_SLASHES + UrlValidator.ALLOW_ALL_SCHEMES, customHostValidator );
3. 整合后的完整代码
import com.linkedin.urls.detection.UrlDetector; import com.linkedin.urls.detection.UrlDetectorOptions; import org.apache.commons.validator.routines.UrlValidator; import org.apache.commons.validator.routines.HostNameValidator; import java.util.ArrayList; import java.util.List; private UrlValidator urlValidator; boolean isValidURL(String url) { if (urlValidator == null) { System.out.println("Creating new URL Validator"); // 自定义主机验证器,允许任意非空主机名 HostNameValidator customHostValidator = host -> host != null && !host.isEmpty(); urlValidator = new UrlValidator( new String[]{"http", "https"}, UrlValidator.ALLOW_2_SLASHES + UrlValidator.ALLOW_ALL_SCHEMES, customHostValidator ); } return urlValidator.isValid(url) || urlValidator.isValid("http://" + url); } // Extract and return valid URL from whole SMS text. private List<String> getURLFromSMSText(String smsTextMessage) { // smsTextMessage = "PARADISE HALEEM @Rs.249 & 1 KG FAMILY PACK @Rs.699. Your Favourite " + // "1 KG PARADISE Nizami Chicken Biryani @Rs.279 " + // "in Hyderabad Valid in Take Away " + // "- http://google/element8441876860/promethium/7893214560.com " + " bit.ly/653282 " + // "sdfsdfsf http://google/googlet.com 7893214560.com" List<String> urlsInSms = new ArrayList<>(); try { // 自定义UrlDetector选项,避免拆分完整URL UrlDetectorOptions options = UrlDetectorOptions.builder() .allowInsecureUrls(true) .allowProtocolRelativeUrls(false) .allowHostlessUrls(false) .build(); UrlDetector parser = new UrlDetector(smsTextMessage, options); List<com.linkedin.urls.Url> found = parser.detect(); if (found != null && !found.isEmpty()) { for (com.linkedin.urls.Url url : found) { String tempUrl = url.getOriginalUrl(); if (tempUrl != null && isValidURL(tempUrl)) { System.out.println("Valid Original URL: " + url.getOriginalUrl()); System.out.println("Scheme: " + url.getScheme()); System.out.println("Host: " + url.getHost()); System.out.println("Path: " + url.getPath()); urlsInSms.add(tempUrl); } else { System.out.println("Retrived Url is In-Valid (" + url.getOriginalUrl() + ")"); if (DebugConfig.getInstance().isDebugEnable()) { log.error("Retrived Url is In-Valid (" + url.getOriginalUrl() + ")"); } } } } else { if (DebugConfig.getInstance().isDebugEnable()) log.error("No Matching URL present in SMS (" + smsTextMessage + ")"); } } catch (Exception ue) { if (DebugConfig.getInstance().isDebugEnable()) log.error("Exception in Url Detection", ue); } return urlsInSms; }
内容的提问来源于stack exchange,提问作者u work
相关产品推荐
相关产品推荐

