Java/GPars线程池疑似堵塞求助:批量检测数据库企业链接有效性
Got it, let's break down how to tackle this problem step by step—whether you're working in Java or Groovy, we can build an efficient workflow to validate those links using curl, plus leverage thread pools to handle parallel checks (critical for keeping things fast when dealing with multiple links per enterprise).
Core Approach
Our goal is threefold: extract hyperlinks from each enterprise's description field, use curl to check each link's HTTP status code (flagging any non-200 responses like 404s), and use thread pools to run these checks in parallel (avoiding the slowdown of sequential processing).
Step-by-Step Implementation
1. Extract Hyperlinks from Description Fields
First, parse the text description to pull out all valid URLs. Regex is the easiest way here—just match standard HTTP/HTTPS patterns:
Groovy Example
def extractUrls(String description) { def urlPattern = /https?:\/\/[^\s]+/ // Matches HTTP/HTTPS URLs until whitespace return description.findAll(urlPattern) } // Usage: def enterpriseDesc = "Check out our site at https://example.com and blog at https://blog.example.com" def urls = extractUrls(enterpriseDesc)
Java Example
import java.util.ArrayList; import java.util.List; import java.util.regex.Matcher; import java.util.regex.Pattern; public class UrlExtractor { public static List<String> extractUrls(String description) { Pattern urlPattern = Pattern.compile("https?://[^\\s]+"); Matcher matcher = urlPattern.matcher(description); List<String> urls = new ArrayList<>(); while (matcher.find()) { urls.add(matcher.group()); } return urls; } // Usage: public static void main(String[] args) { String enterpriseDesc = "Check out our site at https://example.com and blog at https://blog.example.com"; List<String> urls = extractUrls(enterpriseDesc); } }
2. Use curl to Check HTTP Status Codes
We'll use a silent curl command that returns only the HTTP status code, avoiding extra output clutter. Add timeout parameters to prevent hanging on unresponsive links:
Shell Command Base
curl -o /dev/null -s -w "%{http_code}" --connect-timeout 5 --max-time 10 <URL>
-o /dev/null: Discards the response body-s: Runs silently (no progress bars)-w "%{http_code}": Prints only the HTTP status code--connect-timeout/--max-time: Prevents long-running requests
Wrapper Methods for Java/Groovy
Groovy
def getStatusCode(String url) { def process = ["curl", "-o", "/dev/null", "-s", "-w", "%{http_code}", "--connect-timeout", "5", "--max-time", "10", url].execute() process.waitFor() return process.text.trim() // Returns the status code as a string }
Java
import java.io.BufferedReader; import java.io.IOException; import java.io.InputStreamReader; public class CurlStatusChecker { public static String getStatusCode(String url) throws IOException, InterruptedException { ProcessBuilder pb = new ProcessBuilder( "curl", "-o", "/dev/null", "-s", "-w", "%{http_code}", "--connect-timeout", "5", "--max-time", "10", url ); Process process = pb.start(); int exitCode = process.waitFor(); // Handle cases where curl fails (e.g., invalid URL) if (exitCode != 0) { return "CURL_ERROR"; } try (BufferedReader reader = new BufferedReader(new InputStreamReader(process.getInputStream()))) { return reader.readLine().trim(); } } }
3. Parallel Processing with Thread Pools
Batch processing dozens/hundreds of enterprises will be slow if we check links one by one. Use thread pools to run multiple curl checks in parallel—Java's ExecutorService is the foundation here, and Groovy works seamlessly with it.
Java Implementation
import java.util.concurrent.ExecutorService; import java.util.concurrent.Executors; import java.util.concurrent.TimeUnit; public class LinkValidationService { public static void main(String[] args) { // Adjust pool size based on your server's resources (start with 5-10 threads) ExecutorService executor = Executors.newFixedThreadPool(8); // Assume you've fetched a list of enterprises from your database List<Enterprise> enterprises = fetchEnterprisesFromDatabase(); for (Enterprise enterprise : enterprises) { List<String> urls = UrlExtractor.extractUrls(enterprise.getDescription()); for (String url : urls) { executor.submit(() -> { try { String statusCode = CurlStatusChecker.getStatusCode(url); if (!"200".equals(statusCode)) { // Log or save the invalid link (e.g., to a database table) System.out.printf("Enterprise ID: %d | URL: %s | Status Code: %s%n", enterprise.getId(), url, statusCode); } } catch (Exception e) { e.printStackTrace(); } }); } } // Shutdown the executor gracefully executor.shutdown(); executor.awaitTermination(1, TimeUnit.HOURS); // Wait for all tasks to finish } // Dummy method to represent fetching enterprises from your database private static List<Enterprise> fetchEnterprisesFromDatabase() { return new ArrayList<>(); } } // Dummy Enterprise class class Enterprise { private long id; private String description; // Getters/setters omitted for brevity }
Groovy Implementation
Groovy can use the same Java ExecutorService, or simplify parallel code with GPars (though it still relies on Java's thread pool under the hood):
@Grab('org.codehaus.gpars:gpars:1.2.1') import groovyx.gpars.ParallelEnhancer import java.util.concurrent.Executors def executor = Executors.newFixedThreadPool(8) def enterprises = fetchEnterprisesFromDatabase() // Your database fetch method enterprises.each { enterprise -> def urls = extractUrls(enterprise.description) urls.each { url -> executor.submit { def statusCode = getStatusCode(url) if (statusCode != "200") { log.info "Invalid link - Enterprise: ${enterprise.name}, URL: $url, Status: $statusCode" } } } } executor.shutdown() executor.awaitTermination(1, TimeUnit.HOURS) // Alternative with GPars for simpler parallel syntax: // ParallelEnhancer.enhanceInstance(urls) // urls.eachParallel { url -> // def statusCode = getStatusCode(url) // if (statusCode != "200") { // // Log invalid link // } // }
Key Notes & Best Practices
- Curl Dependency: Ensure your runtime environment has
curlinstalled. If you can't use curl, replace it with a pure Java HTTP client like OkHttp or Apache HttpClient (but the thread pool logic stays the same). - Thread Pool Tuning: Adjust the pool size based on your server's CPU and network capacity—too many threads can cause resource exhaustion or get your IP blocked by target websites.
- Error Handling: Add retry logic for transient errors (e.g., 5xx status codes) if needed, and handle cases where
curlfails to execute (e.g., malformed URLs). - Database Batch Processing: Fetch enterprises in small batches instead of loading all records into memory at once to avoid out-of-memory issues.
内容的提问来源于stack exchange,提问作者mike rodent

