如何修改Jena中每条路由的连接数?解决查询超5次线程锁定问题
Hey there, let's dig into this problem you're facing with Jena's QueryEngineHTTP locking up after 5 recursive queries. The issue here is that the addParam("max connections", 10) approach isn't doing what you think it is—those params are for passing SPARQL query parameters, not configuring the underlying HTTP connection pool. Let's fix this step by step:
Jena uses Apache HttpClient under the hood for remote SPARQL queries, and connection limits are controlled by HttpClient's connection manager, not query params. You need to explicitly configure a connection pool with higher per-route and total connection limits.
Here's how to build a custom HttpClient with proper pooling:
import org.apache.http.impl.client.CloseableHttpClient; import org.apache.http.impl.client.HttpClients; import org.apache.http.impl.conn.PoolingHttpClientConnectionManager; import java.util.concurrent.TimeUnit; // Create a connection manager with custom limits PoolingHttpClientConnectionManager connManager = new PoolingHttpClientConnectionManager(); // Total max connections across all routes connManager.setMaxTotal(20); // Max connections allowed for your specific SPARQL endpoint (this is what you need) connManager.setDefaultMaxPerRoute(10); // Build the HttpClient with this manager CloseableHttpClient httpClient = HttpClients.custom() .setConnectionManager(connManager) .build();
Instead of using the default QueryExecutionFactory methods, use the overload that accepts your custom HttpClient to ensure all queries use the pooled connections:
// Inside your recursive function String sparqlEndpoint = "http://your-sparql-endpoint-url"; Query query = QueryFactory.create("YOUR RECURSIVE SPARQL QUERY HERE"); // Use try-with-resources to auto-close QueryExecution and avoid leaks try (QueryExecution qexec = QueryExecutionFactory.create(query, sparqlEndpoint, httpClient)) { ResultSet results = qexec.execSelect(); while (results.hasNext()) { QuerySolution soln = results.nextSolution(); // Pass the same httpClient to your recursive call fetchNarrowerNodes(soln.getResource("targetNode"), httpClient); } } catch (Exception e) { // Handle exceptions to prevent unclosed connections e.printStackTrace(); }
The try-with-resources block above is non-negotiable here. Failing to close QueryExecution instances after use causes connection leaks, which will quickly exhaust your pool and lock up threads—this is likely the main cause of your 5-recursion lockup.
To avoid threads getting stuck indefinitely on slow or unresponsive queries, add timeout configurations to your HttpClient:
import org.apache.http.client.config.RequestConfig; CloseableHttpClient httpClient = HttpClients.custom() .setConnectionManager(connManager) .setConnectionTimeToLive(30, TimeUnit.SECONDS) .setDefaultRequestConfig(RequestConfig.custom() .setConnectTimeout(5000) // 5-second timeout to establish connection .setSocketTimeout(10000) // 10-second timeout for data transfer .build()) .build();
If your SPARQL endpoint supports it, rewrite your logic into a single query using property paths to eliminate client-side recursion entirely. This reduces HTTP calls and avoids connection pool issues altogether:
PREFIX skos: <http://www.w3.org/2004/02/skos/core#> SELECT ?childNode WHERE { ?rootNode skos:narrower+ ?childNode . # Replace with your starting node URI VALUES ?rootNode { <http://your-vocab.com/root-term> } }
That should resolve the thread locking issue. The key mistake was trying to set connection limits via query parameters—those are for endpoint-specific params (like custom timeouts), not the underlying HTTP connection pool.
内容的提问来源于stack exchange,提问作者user8809260

