如何通过PHP实现PDF在Solr中的自动索引与文件名检索?
Hey there! Let's get your PHP + Solr PDF indexing and search workflow sorted out. I’ve tackled this exact setup before, so here’s a step-by-step guide that should solve your issues:
Solr relies on Apache Tika to extract content from PDFs, so first we need to make sure that’s configured properly:
- Enable the Extract Request Handler: Open your Solr core’s
solrconfig.xmland add (or verify) this configuration block:
This maps the extracted PDF content to a<requestHandler name="/update/extract" class="org.apache.solr.handler.extraction.ExtractingRequestHandler"> <lst name="defaults"> <str name="lowernames">true</str> <str name="uprefix">ignored_</str> <str name="fmap.content">content</str> <str name="fmap.filename">filename</str> </lst> </requestHandler>contentfield and the filename to afilenamefield in your Solr schema. - Configure Your Schema: In your core’s
managed-schema(orschema.xml), add these fields if they don’t exist:
The<field name="id" type="string" indexed="true" stored="true" required="true" multiValued="false"/> <field name="filename" type="text_general" indexed="true" stored="true"/> <field name="content" type="text_general" indexed="true" stored="true"/>text_generaltype lets you search both full content and filenames with partial matches; if you want exact filename searches only, usestringinstead. - Check Tika JARs: Solr 7.x and above usually include Tika by default, but if you’re on an older version, make sure the Apache Tika JARs are in your Solr
libdirectory.
Next, we’ll handle the PHP upload and send the PDF directly to Solr for indexing. I recommend using the Solarium library (a popular Solr PHP client) for this—it’s way easier than raw cURL calls.
First, Install Solarium
Use Composer to install it:
composer require solarium/solarium
PHP Upload & Index Code Example
Here’s a complete snippet that handles file uploads, validates PDFs, and indexes them to Solr:
<?php require 'vendor/autoload.php'; // Initialize Solarium client $client = new Solarium\Client([ 'endpoint' => [ 'localhost' => [ 'host' => '127.0.0.1', 'port' => 8983, 'path' => '/solr/', 'core' => 'your_core_name' // Replace with your Solr core name ] ] ]); // Handle file upload if ($_SERVER['REQUEST_METHOD'] === 'POST' && isset($_FILES['pdf_file'])) { $file = $_FILES['pdf_file']; // Validate it's a PDF if ($file['type'] !== 'application/pdf' || pathinfo($file['name'], PATHINFO_EXTENSION) !== 'pdf') { die('Please upload a valid PDF file.'); } // Create an update query with extract handler $update = $client->createUpdate(); $extractQuery = $update->createExtract(); // Set the file to extract $extractQuery->setFile($file['tmp_name']); // Map filename to our Solr field $extractQuery->setParam('resource.name', $file['name']); // Don't just extract—index the content $extractQuery->setParam('extractOnly', 'false'); // Auto-commit the index so changes are immediately available $extractQuery->setParam('commit', 'true'); // Add the extract query to the update and execute $update->addExtract($extractQuery); $result = $client->update($update); echo "PDF uploaded and indexed successfully!"; } ?> <!-- HTML Upload Form --> <form method="POST" enctype="multipart/form-data"> <input type="file" name="pdf_file" accept=".pdf" required> <button type="submit">Upload PDF</button> </form>
If you don’t want to use Solarium, you can use raw cURL instead—here’s a quick snippet for that:
<?php if ($_SERVER['REQUEST_METHOD'] === 'POST' && isset($_FILES['pdf_file'])) { $file = $_FILES['pdf_file']; $solrUrl = 'http://127.0.0.1:8983/solr/your_core_name/update/extract?commit=true&resource.name=' . urlencode($file['name']); $ch = curl_init($solrUrl); curl_setopt($ch, CURLOPT_POST, true); curl_setopt($ch, CURLOPT_RETURNTRANSFER, true); curl_setopt($ch, CURLOPT_POSTFIELDS, [ 'myfile' => new CURLFile($file['tmp_name'], $file['type'], $file['name']) ]); $response = curl_exec($ch); curl_close($ch); echo "PDF indexed: " . $response; } ?>
Now let’s build the search functionality that lets users query text and get back matching PDF filenames. Again, using Solarium makes this straightforward:
<?php require 'vendor/autoload.php'; // Initialize Solarium client (same as before) $client = new Solarium\Client([ 'endpoint' => [ 'localhost' => [ 'host' => '127.0.0.1', 'port' => 8983, 'path' => '/solr/', 'core' => 'your_core_name' ] ] ]); // Handle search request $searchQuery = $_GET['q'] ?? ''; if (!empty($searchQuery)) { // Create a select query $select = $client->createSelect(); $query = $select->createQuery(); // Search both content and filename fields $query->setQuery("content:$searchQuery OR filename:$searchQuery"); // Only return the filename field (plus id if you need it) $select->setFields(['id', 'filename']); // Execute the query $resultset = $client->select($select); // Display results echo "Found " . $resultset->getNumFound() . " results:\n"; foreach ($resultset as $document) { echo "- " . $document->filename . "\n"; } } ?> <!-- HTML Search Form --> <form method="GET"> <input type="text" name="q" placeholder="Search PDF content or filename..." required> <button type="submit">Search</button> </form>
- PDF content not indexing? Double-check that your Solr core’s
solrconfig.xmlhas the Extract Request Handler enabled, and that Tika is properly set up. Also, make sure you’re settingextractOnly=falsein your index request. - Filename search not working? If you used
stringtype for thefilenamefield, you’ll need exact matches. Switch totext_generalif you want partial matches (like searching for "report" and getting "annual_report.pdf"). - Changes not showing up in search? Always include
commit=truein your index request, or set up auto-commit in Solr’ssolrconfig.xml(not recommended for production, but fine for testing).
内容的提问来源于stack exchange,提问作者amit rawat

