如何在Rails应用中正确向FSCrawler上传文档至Elasticsearch索引
搞定Rails上传文件到FSCrawler:完美复刻curl的效果
嘿,我刚好碰到过类似的问题!核心原因是你没摸透curl的-F参数到底在背后做了什么——它自动帮你处理了multipart/form-data请求的所有细节,而Ruby的HTTP客户端需要你手动把这些细节补全。咱们一步步来解决:
首先先纠正你的curl命令(顺便解决临时文件名的问题),再对应到Ruby的实现:
先把curl的问题解决:上传原始文件名
你之前的curl用了临时文件的路径,所以FSCrawler拿到的是临时文件名。只要在-F参数里加上filename=指定原始文件名就行,先验证这个:
curl -F "file=@#{params[:document][:upload].tempfile.path};filename=#{params[:document][:upload].original_filename}" "http://127.0.0.1:8080/fscrawler/_upload?debug=true"
这样上传后,索引里的文件名就会是用户上传的原始名称了,接下来用Ruby客户端复刻这个请求。
推荐方案:用RestClient实现(简洁不易错)
RestClient处理multipart表单比Net::HTTP省心太多,你之前踩的坑都是因为没正确构造表单字段。直接上代码:
require 'rest-client' uploaded_file = params[:document][:upload] response = RestClient.post( 'http://127.0.0.1:8080/fscrawler/_upload?debug=true', { # 关键:这里要把文件包装成带文件名和类型的哈希 file: { filename: uploaded_file.original_filename, content: uploaded_file.tempfile, content_type: uploaded_file.content_type } } ) # 查看响应结果 puts response.body
为啥之前的RestClient尝试失败?
- 直接传临时文件路径:RestClient会把路径当成普通字符串内容发出去,FSCrawler收到的是文本(比如
/tmp/abc123.tmp),自然报Unsupported media type - 用
File.read读内容:这种方式相当于把文件内容当成普通表单字段,没有告诉FSCrawler这是个文件,所以它没法处理,直接返回内部错误
如果必须用Net::HTTP(原生库方案)
Net::HTTP需要手动拼multipart请求的所有部分,比较繁琐,但能精准控制。代码如下:
require 'net/http' require 'uri' require 'cgi' require 'securerandom' uploaded_file = params[:document][:upload] uri = URI('http://127.0.0.1:8080/fscrawler/_upload?debug=true') # 生成随机的请求边界(multipart表单必须的) boundary = "----RubyMultipart#{SecureRandom.hex(16)}" request = Net::HTTP::Post.new(uri) request.content_type = "multipart/form-data; boundary=#{boundary}" # 手动构造请求体 body_parts = [] # 添加file字段的头部信息 body_parts << "--#{boundary}" body_parts << "Content-Disposition: form-data; name=\"file\"; filename=\"#{CGI.escape(uploaded_file.original_filename)}\"" body_parts << "Content-Type: #{uploaded_file.content_type}" body_parts << "" # 读取临时文件的内容 uploaded_file.tempfile.open body_parts << uploaded_file.tempfile.read uploaded_file.tempfile.close # 结束请求边界 body_parts << "--#{boundary}--" body_parts << "" request.body = body_parts.join("\r\n") # 发送请求 response = Net::HTTP.start(uri.hostname, uri.port) do |http| http.request(request) end puts response.body
之前Net::HTTP踩坑的原因:
- 直接传Tempfile对象:没设置
filename和正确的Content-Disposition,FSCrawler根本识别不出这是个有效文件,所以Tika报了ZeroByteFileException(其实是请求格式不对,不是文件真的为空) - 传tempfile.path:Net::HTTP把路径当成字符串发了,FSCrawler收到的是路径文本,所以索引的content字段就是这个路径,而不是文件内容
最后验证两个关键点
- 一定要确保请求的Content-Type是
multipart/form-data,并且带正确的边界值 file字段的Content-Disposition里必须包含filename=,指定原始文件名
按这个来,就能完全复刻curl的效果,既上传正确的文件内容,又保留原始文件名啦!
内容的提问来源于stack exchange,提问作者David Krider
相关产品推荐
相关产品推荐

