如何让Solarium PHP用字符串替代文件调用file_get_contents上传PDF至Solr
解决方案:用Solarium索引已下载为字符串的PDF
方法1:使用内存流(推荐)
利用PHP的php://memory或php://temp内存流,将PDF字符串转为可读流资源,直接传递给Solarium,无需临时文件。
代码示例
// 假设$pdfContent是爬虫已下载的PDF字符串 $stream = fopen('php://memory', 'r+'); fwrite($stream, $pdfContent); rewind($stream); // 将指针重置到流开头,确保Solarium能读取完整内容 // 构建Solarium文档 $doc = $client->createDocument(); $doc->id = 'unique-doc-id-123'; // 替换为你的文档唯一ID $doc->file = $stream; // 对应Solr配置中用于提取内容的字段(需提前配置Solr的提取处理器) // 提交索引请求 $update = $client->createUpdate(); $update->addDocuments([$doc]); $update->addCommit(); $client->update($update); // 操作完成后关闭流 fclose($stream);
原理
Solarium的buildUploadBodyFromRequest方法中,file_get_contents不仅能读取文件路径,也能处理有效的流资源。内存流完全在内存中操作,不会产生磁盘IO,也不影响爬虫的并行能力。
方法2:自定义AdapterHelper扩展Solarium逻辑
如果内存流方案不符合需求,可以重写Solarium的AdapterHelper类,直接处理字符串类型的文件内容,避免file_get_contents的文件依赖。
步骤1:创建自定义Helper类
use Solarium\Core\Client\Adapter\AdapterHelper; class StringAwareAdapterHelper extends AdapterHelper { protected function buildUploadBodyFromRequest($request) { $body = ''; $boundary = $this->getMultipartBoundary($request); foreach ($request->getFiles() as $file) { $body .= "--{$boundary}\r\n"; $body .= 'Content-Disposition: form-data; name="' . $file['name'] . '"; filename="' . $file['filename'] . '"' . "\r\n"; $body .= 'Content-Type: ' . ($file['type'] ?? 'application/pdf') . "\r\n\r\n"; // 直接使用字符串内容,无需读取文件 if (is_string($file['content'])) { $body .= $file['content'] . "\r\n"; } else { // 回退到原逻辑处理文件/流 $body .= file_get_contents($file['content']) . "\r\n"; } } $body .= "--{$boundary}--\r\n"; return $body; } }
步骤2:替换Solarium的AdapterHelper
初始化客户端时,将自定义Helper绑定到适配器:
// 创建Curl适配器并替换Helper $adapter = new \Solarium\Core\Client\Adapter\Curl(); $adapter->setHelper(new StringAwareAdapterHelper()); // 初始化Solarium客户端 $client = new \Solarium\Client($adapter, new \Solarium\Core\Client\Endpoint([ 'host' => 'your-solr-host', 'port' => 8983, 'path' => '/solr/your-core-name', ]));
使用方式
构建文档时直接传入PDF字符串:
$doc = $client->createDocument(); $doc->id = 'unique-doc-id-456'; $doc->file = [ 'content' => $pdfContent, 'name' => 'file', // 对应Solr请求的参数名 'filename' => 'document.pdf', // 用于Solr识别文件类型的文件名 'type' => 'application/pdf' ]; // 提交索引请求(同方法1)
方法3:直接构造Multipart POST请求(绕过Solarium)
如果不想依赖Solarium的文档处理逻辑,可以利用你已有的curl_multi_*爬虫,直接构造符合Solr要求的Multipart请求上传PDF字符串,保持并行下载的优势。这种灵活性最高,但需要手动处理Solr的请求格式。
内容的提问来源于stack exchange,提问作者rexkogitans
相关产品推荐
相关产品推荐

