You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过PHP实现PDF在Solr中的自动索引与文件名检索?

Hey there! Let's get your PHP + Solr PDF indexing and search workflow sorted out. I’ve tackled this exact setup before, so here’s a step-by-step guide that should solve your issues:

Step 1: Get Solr Ready to Parse PDFs

Solr relies on Apache Tika to extract content from PDFs, so first we need to make sure that’s configured properly:

  • Enable the Extract Request Handler: Open your Solr core’s solrconfig.xml and add (or verify) this configuration block:
    <requestHandler name="/update/extract" class="org.apache.solr.handler.extraction.ExtractingRequestHandler">
      <lst name="defaults">
        <str name="lowernames">true</str>
        <str name="uprefix">ignored_</str>
        <str name="fmap.content">content</str>
        <str name="fmap.filename">filename</str>
      </lst>
    </requestHandler>
    
    This maps the extracted PDF content to a content field and the filename to a filename field in your Solr schema.
  • Configure Your Schema: In your core’s managed-schema (or schema.xml), add these fields if they don’t exist:
    <field name="id" type="string" indexed="true" stored="true" required="true" multiValued="false"/>
    <field name="filename" type="text_general" indexed="true" stored="true"/>
    <field name="content" type="text_general" indexed="true" stored="true"/>
    
    The text_general type lets you search both full content and filenames with partial matches; if you want exact filename searches only, use string instead.
  • Check Tika JARs: Solr 7.x and above usually include Tika by default, but if you’re on an older version, make sure the Apache Tika JARs are in your Solr lib directory.
Step 2: PHP Upload + Auto-Index to Solr

Next, we’ll handle the PHP upload and send the PDF directly to Solr for indexing. I recommend using the Solarium library (a popular Solr PHP client) for this—it’s way easier than raw cURL calls.

First, Install Solarium

Use Composer to install it:

composer require solarium/solarium

PHP Upload & Index Code Example

Here’s a complete snippet that handles file uploads, validates PDFs, and indexes them to Solr:

<?php
require 'vendor/autoload.php';

// Initialize Solarium client
$client = new Solarium\Client([
    'endpoint' => [
        'localhost' => [
            'host' => '127.0.0.1',
            'port' => 8983,
            'path' => '/solr/',
            'core' => 'your_core_name' // Replace with your Solr core name
        ]
    ]
]);

// Handle file upload
if ($_SERVER['REQUEST_METHOD'] === 'POST' && isset($_FILES['pdf_file'])) {
    $file = $_FILES['pdf_file'];
    
    // Validate it's a PDF
    if ($file['type'] !== 'application/pdf' || pathinfo($file['name'], PATHINFO_EXTENSION) !== 'pdf') {
        die('Please upload a valid PDF file.');
    }
    
    // Create an update query with extract handler
    $update = $client->createUpdate();
    $extractQuery = $update->createExtract();
    
    // Set the file to extract
    $extractQuery->setFile($file['tmp_name']);
    // Map filename to our Solr field
    $extractQuery->setParam('resource.name', $file['name']);
    // Don't just extract—index the content
    $extractQuery->setParam('extractOnly', 'false');
    // Auto-commit the index so changes are immediately available
    $extractQuery->setParam('commit', 'true');
    
    // Add the extract query to the update and execute
    $update->addExtract($extractQuery);
    $result = $client->update($update);
    
    echo "PDF uploaded and indexed successfully!";
}
?>

<!-- HTML Upload Form -->
<form method="POST" enctype="multipart/form-data">
    <input type="file" name="pdf_file" accept=".pdf" required>
    <button type="submit">Upload PDF</button>
</form>

If you don’t want to use Solarium, you can use raw cURL instead—here’s a quick snippet for that:

<?php
if ($_SERVER['REQUEST_METHOD'] === 'POST' && isset($_FILES['pdf_file'])) {
    $file = $_FILES['pdf_file'];
    $solrUrl = 'http://127.0.0.1:8983/solr/your_core_name/update/extract?commit=true&resource.name=' . urlencode($file['name']);
    
    $ch = curl_init($solrUrl);
    curl_setopt($ch, CURLOPT_POST, true);
    curl_setopt($ch, CURLOPT_RETURNTRANSFER, true);
    curl_setopt($ch, CURLOPT_POSTFIELDS, [
        'myfile' => new CURLFile($file['tmp_name'], $file['type'], $file['name'])
    ]);
    
    $response = curl_exec($ch);
    curl_close($ch);
    
    echo "PDF indexed: " . $response;
}
?>
Step 3: PHP Search to Return PDF Filenames

Now let’s build the search functionality that lets users query text and get back matching PDF filenames. Again, using Solarium makes this straightforward:

<?php
require 'vendor/autoload.php';

// Initialize Solarium client (same as before)
$client = new Solarium\Client([
    'endpoint' => [
        'localhost' => [
            'host' => '127.0.0.1',
            'port' => 8983,
            'path' => '/solr/',
            'core' => 'your_core_name'
        ]
    ]
]);

// Handle search request
$searchQuery = $_GET['q'] ?? '';
if (!empty($searchQuery)) {
    // Create a select query
    $select = $client->createSelect();
    $query = $select->createQuery();
    
    // Search both content and filename fields
    $query->setQuery("content:$searchQuery OR filename:$searchQuery");
    // Only return the filename field (plus id if you need it)
    $select->setFields(['id', 'filename']);
    
    // Execute the query
    $resultset = $client->select($select);
    
    // Display results
    echo "Found " . $resultset->getNumFound() . " results:\n";
    foreach ($resultset as $document) {
        echo "- " . $document->filename . "\n";
    }
}
?>

<!-- HTML Search Form -->
<form method="GET">
    <input type="text" name="q" placeholder="Search PDF content or filename..." required>
    <button type="submit">Search</button>
</form>
Troubleshooting Common Issues
  • PDF content not indexing? Double-check that your Solr core’s solrconfig.xml has the Extract Request Handler enabled, and that Tika is properly set up. Also, make sure you’re setting extractOnly=false in your index request.
  • Filename search not working? If you used string type for the filename field, you’ll need exact matches. Switch to text_general if you want partial matches (like searching for "report" and getting "annual_report.pdf").
  • Changes not showing up in search? Always include commit=true in your index request, or set up auto-commit in Solr’s solrconfig.xml (not recommended for production, but fine for testing).

内容的提问来源于stack exchange,提问作者amit rawat

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 06:52:52