Ruby on Rails:如何将Word文档中的图表等元素转换为图片
解决方案:将Word文档中的图表/图形转为图片(适配你的Rails技术栈)
我之前处理过类似的Office文档图表渲染问题,结合你使用的Rails 5.0.7、Ruby 2.5.7和Selenium 3.x技术栈,这里提供一个可行的方案——核心思路是先把docx转成可被浏览器解析的HTML,用Headless Selenium截图所有图表元素,再把截图替换回原docx的对应位置,这样上传到Google Drive后转HTML时就会显示正常的图片了。
前置准备
首先需要安装几个工具和Gem:
- 服务器上安装LibreOffice(确保docx转HTML的兼容性,Ubuntu命令:
sudo apt install libreoffice) - 在Gemfile中添加依赖:
gem 'mammoth' # 转docx到HTML gem 'selenium-webdriver', '3.142.7' gem 'mini_magick' # 图片处理 gem 'rubyzip' # 解压/打包docx gem 'nokogiri'
然后运行bundle install。
具体实现步骤
1. 将上传的docx转为HTML
用mammoth把docx转成HTML,它会保留原文档中drawing元素对应的结构,生成带有特定类名的元素(比如.mammoth-drawing),方便后续Selenium定位。
require 'mammoth' require 'rubyzip' require 'selenium-webdriver' require 'mini_magick' require 'base64' def convert_docx_to_html(docx_path) result = Mammoth.convert_to_html(docx_path) html_content = result.value # 生成唯一的临时HTML文件 html_path = "#{Rails.root}/tmp/converted_#{SecureRandom.hex(8)}.html" File.write(html_path, html_content) html_path end
2. 用Headless Selenium截图所有图表元素
这里解决了你之前直接打开docx的问题——先转成HTML,浏览器可以正常解析。定位.mammoth-drawing元素并截图,保存为Base64格式方便后续替换。
def capture_drawing_elements(html_path) # 配置Firefox Headless模式 options = Selenium::WebDriver::Firefox::Options.new(binary: '/usr/bin/firefox', headless: true) driver = Selenium::WebDriver.for :firefox, options: options # 正确处理本地文件路径(转义特殊字符) encoded_path = ERB::Util.url_encode(html_path) driver.navigate.to("file://#{encoded_path}") # 定位所有drawing对应的HTML元素 drawing_elements = driver.find_elements(:css, '.mammoth-drawing') screenshots = [] drawing_elements.each_with_index do |element, index| img_path = "#{Rails.root}/tmp/drawing_#{index}.png" element.screenshot.save(img_path) # 转成Base64避免文件IO问题 img_base64 = Base64.strict_encode64(File.read(img_path)) screenshots << img_base64 File.delete(img_path) # 清理临时文件 end driver.quit screenshots end
3. 将截图替换回原docx的drawing节点
解压docx包,修改word/document.xml,把所有<w:drawing>节点替换为标准的Word图片节点,同时更新文档的关系文件(确保图片能被正确引用)。
def replace_drawings_with_images(docx_path, screenshots) tmp_dir = "#{Rails.root}/tmp/docx_tmp_#{SecureRandom.hex(8)}" FileUtils.mkdir_p(tmp_dir) # 解压原docx到临时目录 Zip::File.open(docx_path) do |zip| zip.each do |entry| entry.extract("#{tmp_dir}/#{entry.name}") end end # 读取并修改document.xml doc_xml_path = "#{tmp_dir}/word/document.xml" doc = Nokogiri::XML(File.read(doc_xml_path)) drawing_nodes = doc.xpath('//w:drawing', 'w' => 'http://schemas.openxmlformats.org/wordprocessingml/2006/main') drawing_nodes.each_with_index do |node, index| next if index >= screenshots.size # 保存图片到docx的media目录 img_filename = "media/image_#{index}.png" img_path = "#{tmp_dir}/word/#{img_filename}" File.write(img_path, Base64.strict_decode64(screenshots[index])) # 构建Word标准的图片XML节点,替换原drawing节点 image_node = build_word_image_node(doc, img_filename, index) node.replace(image_node) end # 更新关系文件,添加图片引用 update_rels_file("#{tmp_dir}/word/_rels/document.xml.rels", screenshots.size) # 重新打包成docx new_docx_path = "#{Rails.root}/tmp/modified_#{SecureRandom.hex(8)}.docx" Zip::File.open(new_docx_path, Zip::File::CREATE) do |zip| Dir.glob("#{tmp_dir}/**/*").each do |file| next if File.directory?(file) relative_path = file.sub("#{tmp_dir}/", '') zip.add(relative_path, file) end end # 清理临时目录 FileUtils.rm_rf(tmp_dir) new_docx_path end # 辅助方法:构建符合Word OOXML标准的图片节点 def build_word_image_node(doc, img_filename, index) ns = { 'w' => 'http://schemas.openxmlformats.org/wordprocessingml/2006/main', 'r' => 'http://schemas.openxmlformats.org/officeDocument/2006/relationships', 'wp' => 'http://schemas.openxmlformats.org/drawingml/2006/wordprocessingDrawing', 'a' => 'http://schemas.openxmlformats.org/drawingml/2006/main', 'pic' => 'http://schemas.openxmlformats.org/drawingml/2006/picture' } # 构建inline图片结构(Word中嵌入图片的标准格式) inline = doc.create_element('wp:inline', ns) inline.set_attributes({ distT: '0', distB: '0', distL: '0', distR: '0' }) extent = doc.create_element('wp:extent', ns) extent.set_attributes({ cx: '914400', cy: '685800' }) # 1x0.75英寸,可根据截图调整 inline.add_child(extent) doc_pr = doc.create_element('wp:docPr', ns) doc_pr.set_attributes({ id: "#{index + 1}", name: "Picture #{index + 1}" }) inline.add_child(doc_pr) graphic = doc.create_element('a:graphic', ns) graphic_data = doc.create_element('a:graphicData', ns) graphic_data.set_attribute('uri', 'http://schemas.openxmlformats.org/drawingml/2006/picture') pic = doc.create_element('pic:pic', ns) nv_pic_pr = doc.create_element('pic:nvPicPr', ns) c_nv_pr = doc.create_element('pic:cNvPr', ns) c_nv_pr.set_attributes({ id: "#{index + 1}", name: "Picture #{index + 1}" }) nv_pic_pr.add_child(c_nv_pr) pic.add_child(nv_pic_pr) blip_fill = doc.create_element('pic:blipFill', ns) blip = doc.create_element('a:blip', ns) blip.set_attribute('r:embed', "rId#{index + 10}") # rId要和rels文件对应 blip_fill.add_child(blip) pic.add_child(blip_fill) graphic_data.add_child(pic) graphic.add_child(graphic_data) inline.add_child(graphic) # 包装在r节点中(Word段落中的元素容器) r = doc.create_element('w:r', ns) r.add_child(doc.create_element('w:rPr', ns)) r.add_child(inline) r end # 辅助方法:更新rels文件,添加图片的关系引用 def update_rels_file(rels_path, count) doc = Nokogiri::XML(File.read(rels_path)) ns = { 'r' => 'http://schemas.openxmlformats.org/package/2006/relationships' } # 获取现有最大rId,避免冲突 existing_rids = doc.xpath('//r:Relationship/@Id', ns).map { |id| id.value.gsub('rId', '').to_i } max_rid = existing_rids.max || 0 count.times do |i| rid = "rId#{max_rid + i + 1}" rel = doc.create_element('r:Relationship', ns) rel.set_attributes({ Id: rid, Type: 'http://schemas.openxmlformats.org/officeDocument/2006/relationships/image', Target: "media/image_#{i}.png" }) doc.root.add_child(rel) end File.write(rels_path, doc.to_xml) end
4. 在上传流程中整合所有方法
# 假设这是你的上传处理方法,接收上传的文件对象 def process_uploaded_docx(uploaded_file) tmp_docx_path = "#{Rails.root}/tmp/uploaded_#{SecureRandom.hex(8)}.docx" File.write(tmp_docx_path, uploaded_file.read) begin html_path = convert_docx_to_html(tmp_docx_path) screenshots = capture_drawing_elements(html_path) modified_docx_path = replace_drawings_with_images(tmp_docx_path, screenshots) # 这里替换成你上传到Google Drive的逻辑 # google_drive_service.upload_file(modified_docx_path, name: 'processed_document.docx') # 清理临时文件 [tmp_docx_path, html_path, modified_docx_path].each do |path| File.delete(path) if File.exist?(path) end rescue => e Rails.logger.error("Docx processing failed: #{e.message}\n#{e.backtrace.join("\n")}") # 异常时清理临时文件 [tmp_docx_path, html_path, modified_docx_path].each do |path| File.delete(path) if File.exist?(path) end raise e # 根据业务需求决定是否抛出异常 end end
关键优势
- 解决Selenium无法打开docx的问题:先转成HTML,浏览器可以正常解析渲染。
- 支持所有类型的drawing元素:不管是Excel图表、SVG还是复杂图形,只要能在HTML中渲染,就能被截图。
- 保留标准docx格式:替换后的文档符合Office Open XML规范,Google Drive可以正常处理,转HTML时会显示图片而非异常的drawing元素。
注意事项
- 确保服务器上安装了Firefox和对应版本的Geckodriver(适配Selenium 3.142.7)。
- 可以根据截图的实际尺寸调整图片节点的
cx和cy属性(1英寸=914400 EMU单位)。 - 处理大型文档时,建议增加Selenium的超时时间,避免截图失败。
内容的提问来源于stack exchange,提问作者altose87
相关产品推荐
相关产品推荐

