You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Java/GPars线程池疑似堵塞求助:批量检测数据库企业链接有效性

解决企业数据表链接有效性检测问题(Java/Groovy 环境适配)

Got it, let's break down how to tackle this problem step by step—whether you're working in Java or Groovy, we can build an efficient workflow to validate those links using curl, plus leverage thread pools to handle parallel checks (critical for keeping things fast when dealing with multiple links per enterprise).

Core Approach

Our goal is threefold: extract hyperlinks from each enterprise's description field, use curl to check each link's HTTP status code (flagging any non-200 responses like 404s), and use thread pools to run these checks in parallel (avoiding the slowdown of sequential processing).

Step-by-Step Implementation

First, parse the text description to pull out all valid URLs. Regex is the easiest way here—just match standard HTTP/HTTPS patterns:

Groovy Example

def extractUrls(String description) {
    def urlPattern = /https?:\/\/[^\s]+/ // Matches HTTP/HTTPS URLs until whitespace
    return description.findAll(urlPattern)
}

// Usage:
def enterpriseDesc = "Check out our site at https://example.com and blog at https://blog.example.com"
def urls = extractUrls(enterpriseDesc)

Java Example

import java.util.ArrayList;
import java.util.List;
import java.util.regex.Matcher;
import java.util.regex.Pattern;

public class UrlExtractor {
    public static List<String> extractUrls(String description) {
        Pattern urlPattern = Pattern.compile("https?://[^\\s]+");
        Matcher matcher = urlPattern.matcher(description);
        List<String> urls = new ArrayList<>();
        while (matcher.find()) {
            urls.add(matcher.group());
        }
        return urls;
    }

    // Usage:
    public static void main(String[] args) {
        String enterpriseDesc = "Check out our site at https://example.com and blog at https://blog.example.com";
        List<String> urls = extractUrls(enterpriseDesc);
    }
}

2. Use curl to Check HTTP Status Codes

We'll use a silent curl command that returns only the HTTP status code, avoiding extra output clutter. Add timeout parameters to prevent hanging on unresponsive links:

Shell Command Base

curl -o /dev/null -s -w "%{http_code}" --connect-timeout 5 --max-time 10 <URL>
  • -o /dev/null: Discards the response body
  • -s: Runs silently (no progress bars)
  • -w "%{http_code}": Prints only the HTTP status code
  • --connect-timeout/--max-time: Prevents long-running requests

Wrapper Methods for Java/Groovy

Groovy
def getStatusCode(String url) {
    def process = ["curl", "-o", "/dev/null", "-s", "-w", "%{http_code}", "--connect-timeout", "5", "--max-time", "10", url].execute()
    process.waitFor()
    return process.text.trim() // Returns the status code as a string
}
Java
import java.io.BufferedReader;
import java.io.IOException;
import java.io.InputStreamReader;

public class CurlStatusChecker {
    public static String getStatusCode(String url) throws IOException, InterruptedException {
        ProcessBuilder pb = new ProcessBuilder(
            "curl", "-o", "/dev/null", "-s", "-w", "%{http_code}",
            "--connect-timeout", "5", "--max-time", "10", url
        );
        Process process = pb.start();
        int exitCode = process.waitFor();
        
        // Handle cases where curl fails (e.g., invalid URL)
        if (exitCode != 0) {
            return "CURL_ERROR";
        }
        
        try (BufferedReader reader = new BufferedReader(new InputStreamReader(process.getInputStream()))) {
            return reader.readLine().trim();
        }
    }
}

3. Parallel Processing with Thread Pools

Batch processing dozens/hundreds of enterprises will be slow if we check links one by one. Use thread pools to run multiple curl checks in parallel—Java's ExecutorService is the foundation here, and Groovy works seamlessly with it.

Java Implementation

import java.util.concurrent.ExecutorService;
import java.util.concurrent.Executors;
import java.util.concurrent.TimeUnit;

public class LinkValidationService {
    public static void main(String[] args) {
        // Adjust pool size based on your server's resources (start with 5-10 threads)
        ExecutorService executor = Executors.newFixedThreadPool(8);
        
        // Assume you've fetched a list of enterprises from your database
        List<Enterprise> enterprises = fetchEnterprisesFromDatabase();
        
        for (Enterprise enterprise : enterprises) {
            List<String> urls = UrlExtractor.extractUrls(enterprise.getDescription());
            for (String url : urls) {
                executor.submit(() -> {
                    try {
                        String statusCode = CurlStatusChecker.getStatusCode(url);
                        if (!"200".equals(statusCode)) {
                            // Log or save the invalid link (e.g., to a database table)
                            System.out.printf("Enterprise ID: %d | URL: %s | Status Code: %s%n",
                                enterprise.getId(), url, statusCode);
                        }
                    } catch (Exception e) {
                        e.printStackTrace();
                    }
                });
            }
        }
        
        // Shutdown the executor gracefully
        executor.shutdown();
        executor.awaitTermination(1, TimeUnit.HOURS); // Wait for all tasks to finish
    }
    
    // Dummy method to represent fetching enterprises from your database
    private static List<Enterprise> fetchEnterprisesFromDatabase() {
        return new ArrayList<>();
    }
}

// Dummy Enterprise class
class Enterprise {
    private long id;
    private String description;
    
    // Getters/setters omitted for brevity
}

Groovy Implementation

Groovy can use the same Java ExecutorService, or simplify parallel code with GPars (though it still relies on Java's thread pool under the hood):

@Grab('org.codehaus.gpars:gpars:1.2.1')
import groovyx.gpars.ParallelEnhancer
import java.util.concurrent.Executors

def executor = Executors.newFixedThreadPool(8)
def enterprises = fetchEnterprisesFromDatabase() // Your database fetch method

enterprises.each { enterprise ->
    def urls = extractUrls(enterprise.description)
    urls.each { url ->
        executor.submit {
            def statusCode = getStatusCode(url)
            if (statusCode != "200") {
                log.info "Invalid link - Enterprise: ${enterprise.name}, URL: $url, Status: $statusCode"
            }
        }
    }
}

executor.shutdown()
executor.awaitTermination(1, TimeUnit.HOURS)

// Alternative with GPars for simpler parallel syntax:
// ParallelEnhancer.enhanceInstance(urls)
// urls.eachParallel { url ->
//     def statusCode = getStatusCode(url)
//     if (statusCode != "200") {
//         // Log invalid link
//     }
// }

Key Notes & Best Practices

  • Curl Dependency: Ensure your runtime environment has curl installed. If you can't use curl, replace it with a pure Java HTTP client like OkHttp or Apache HttpClient (but the thread pool logic stays the same).
  • Thread Pool Tuning: Adjust the pool size based on your server's CPU and network capacity—too many threads can cause resource exhaustion or get your IP blocked by target websites.
  • Error Handling: Add retry logic for transient errors (e.g., 5xx status codes) if needed, and handle cases where curl fails to execute (e.g., malformed URLs).
  • Database Batch Processing: Fetch enterprises in small batches instead of loading all records into memory at once to avoid out-of-memory issues.

内容的提问来源于stack exchange,提问作者mike rodent

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:50:28