Apache Tika 1.18在SpringBoot中BodyContentHandler返回空值求助
Hey there, let's work through this frustrating issue where Apache Tika 1.18 returns empty output in Spring Boot but works fine in SparkJava. Here are some targeted troubleshooting steps to fix this:
Check Thread Safety of Tika Components
Apache Tika's BodyContentHandler and AutoDetectParser are not thread-safe. In SparkJava, you might be creating new instances of these classes for every request, which avoids concurrency issues. But if you've defined them as singleton beans in Spring Boot (via @Bean without specifying a scope), concurrent requests will corrupt the handlers and lead to empty outputs.
Fix this by either:
- Creating new Tika instances per request, or
- Setting the bean scope to
prototypeso Spring creates a fresh instance each time it's injected:
@Bean @Scope("prototype") public BodyContentHandler bodyContentHandler() { return new BodyContentHandler(); } @Bean @Scope("prototype") public AutoDetectParser autoDetectParser() { return new AutoDetectParser(); }
Verify Stream Handling After Decoding
Even if your decoded byte array matches what works in SparkJava, make sure the input stream passed to Tika is read from the start. Spring Boot's request processing might accidentally leave streams in an EOF state if they're read elsewhere before Tika gets to them.
When using your decoded bytes, wrap them in a fresh ByteArrayInputStream and avoid any pre-reading:
// After URL + Base64 decoding byte[] decodedContent = ...; try (InputStream tikaInput = new ByteArrayInputStream(decodedContent)) { BodyContentHandler handler = new BodyContentHandler(); Metadata metadata = new Metadata(); AutoDetectParser parser = new AutoDetectParser(); parser.parse(tikaInput, handler, metadata, new ParseContext()); String extractedText = handler.toString(); } catch (IOException | TikaException e) { // Handle exceptions properly }
Inspect Spring Boot Request Filters
Some Spring Boot filters (like multipart processing filters or custom logging filters) read the request body stream before your controller gets it. By the time Tika tries to parse the stream, it's already exhausted.
To bypass this, read the full request body into a byte array early in your controller:
@PostMapping("/extract-text") public ResponseEntity<String> extractText(HttpServletRequest request) throws IOException { // Read the entire request body first byte[] rawRequest = IOUtils.toByteArray(request.getInputStream()); // Perform your URL and Base64 decoding String urlDecoded = URLDecoder.decode(new String(rawRequest, StandardCharsets.UTF_8), StandardCharsets.UTF_8.name()); byte[] base64Decoded = Base64.getDecoder().decode(urlDecoded); // Now pass the decoded bytes to Tika try (InputStream inputStream = new ByteArrayInputStream(base64Decoded)) { BodyContentHandler handler = new BodyContentHandler(); new AutoDetectParser().parse(inputStream, handler, new Metadata(), new ParseContext()); return ResponseEntity.ok(handler.toString()); } catch (TikaException e) { return ResponseEntity.badRequest().body("Failed to parse content: " + e.getMessage()); } }
Check for Dependency Conflicts
Spring Boot's dependency management might pull in newer versions of libraries that Tika 1.18 doesn't play well with (like commons-io or slf4j). Use your build tool to check for conflicts:
- For Maven: Run
mvn dependency:treeand look for mismatched versions - For Gradle: Run
./gradlew dependencies
If you find conflicting dependencies, force the version required by Tika 1.18. For example, in Maven:
<dependency> <groupId>commons-io</groupId> <artifactId>commons-io</artifactId> <version>2.6</version> <!-- Matches Tika 1.18's dependency --> </dependency>
Enable Tika Debug Logs
Turn on debug logging for Tika to see exactly what's happening during parsing. Add this to your application.properties:
logging.level.org.apache.tika=DEBUG
You'll get detailed logs about file type detection, parsing steps, and any hidden exceptions that might be causing the empty output.
内容的提问来源于stack exchange,提问作者Morkus

