You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何让Solarium PHP用字符串替代文件调用file_get_contents上传PDF至Solr

解决方案:用Solarium索引已下载为字符串的PDF

方法1:使用内存流(推荐)

利用PHP的php://memory或php://temp内存流,将PDF字符串转为可读流资源,直接传递给Solarium,无需临时文件。

代码示例

// 假设$pdfContent是爬虫已下载的PDF字符串
$stream = fopen('php://memory', 'r+');
fwrite($stream, $pdfContent);
rewind($stream); // 将指针重置到流开头,确保Solarium能读取完整内容

// 构建Solarium文档
$doc = $client->createDocument();
$doc->id = 'unique-doc-id-123'; // 替换为你的文档唯一ID
$doc->file = $stream; // 对应Solr配置中用于提取内容的字段(需提前配置Solr的提取处理器)

// 提交索引请求
$update = $client->createUpdate();
$update->addDocuments([$doc]);
$update->addCommit();
$client->update($update);

// 操作完成后关闭流
fclose($stream);

原理

Solarium的buildUploadBodyFromRequest方法中,file_get_contents不仅能读取文件路径,也能处理有效的流资源。内存流完全在内存中操作,不会产生磁盘IO,也不影响爬虫的并行能力。


方法2:自定义AdapterHelper扩展Solarium逻辑

如果内存流方案不符合需求,可以重写Solarium的AdapterHelper类,直接处理字符串类型的文件内容,避免file_get_contents的文件依赖。

步骤1:创建自定义Helper类

use Solarium\Core\Client\Adapter\AdapterHelper;

class StringAwareAdapterHelper extends AdapterHelper
{
    protected function buildUploadBodyFromRequest($request)
    {
        $body = '';
        $boundary = $this->getMultipartBoundary($request);

        foreach ($request->getFiles() as $file) {
            $body .= "--{$boundary}\r\n";
            $body .= 'Content-Disposition: form-data; name="' . $file['name'] . '"; filename="' . $file['filename'] . '"' . "\r\n";
            $body .= 'Content-Type: ' . ($file['type'] ?? 'application/pdf') . "\r\n\r\n";

            // 直接使用字符串内容,无需读取文件
            if (is_string($file['content'])) {
                $body .= $file['content'] . "\r\n";
            } else {
                // 回退到原逻辑处理文件/流
                $body .= file_get_contents($file['content']) . "\r\n";
            }
        }

        $body .= "--{$boundary}--\r\n";
        return $body;
    }
}

步骤2:替换Solarium的AdapterHelper

初始化客户端时,将自定义Helper绑定到适配器:

// 创建Curl适配器并替换Helper
$adapter = new \Solarium\Core\Client\Adapter\Curl();
$adapter->setHelper(new StringAwareAdapterHelper());

// 初始化Solarium客户端
$client = new \Solarium\Client($adapter, new \Solarium\Core\Client\Endpoint([
    'host' => 'your-solr-host',
    'port' => 8983,
    'path' => '/solr/your-core-name',
]));

使用方式

构建文档时直接传入PDF字符串:

$doc = $client->createDocument();
$doc->id = 'unique-doc-id-456';
$doc->file = [
    'content' => $pdfContent,
    'name' => 'file', // 对应Solr请求的参数名
    'filename' => 'document.pdf', // 用于Solr识别文件类型的文件名
    'type' => 'application/pdf'
];

// 提交索引请求(同方法1)

方法3:直接构造Multipart POST请求(绕过Solarium)

如果不想依赖Solarium的文档处理逻辑,可以利用你已有的curl_multi_*爬虫,直接构造符合Solr要求的Multipart请求上传PDF字符串,保持并行下载的优势。这种灵活性最高,但需要手动处理Solr的请求格式。

内容的提问来源于stack exchange,提问作者rexkogitans

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 20:03:30