You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Ruby on Rails:如何将Word文档中的图表等元素转换为图片

解决方案:将Word文档中的图表/图形转为图片(适配你的Rails技术栈)

我之前处理过类似的Office文档图表渲染问题,结合你使用的Rails 5.0.7、Ruby 2.5.7和Selenium 3.x技术栈,这里提供一个可行的方案——核心思路是先把docx转成可被浏览器解析的HTML,用Headless Selenium截图所有图表元素,再把截图替换回原docx的对应位置,这样上传到Google Drive后转HTML时就会显示正常的图片了。

前置准备

首先需要安装几个工具和Gem:

  1. 服务器上安装LibreOffice(确保docx转HTML的兼容性,Ubuntu命令:sudo apt install libreoffice)
  2. 在Gemfile中添加依赖:
gem 'mammoth'       # 转docx到HTML
gem 'selenium-webdriver', '3.142.7'
gem 'mini_magick'   # 图片处理
gem 'rubyzip'       # 解压/打包docx
gem 'nokogiri'

然后运行bundle install。

具体实现步骤

1. 将上传的docx转为HTML

用mammoth把docx转成HTML,它会保留原文档中drawing元素对应的结构,生成带有特定类名的元素(比如.mammoth-drawing),方便后续Selenium定位。

require 'mammoth'
require 'rubyzip'
require 'selenium-webdriver'
require 'mini_magick'
require 'base64'

def convert_docx_to_html(docx_path)
  result = Mammoth.convert_to_html(docx_path)
  html_content = result.value
  # 生成唯一的临时HTML文件
  html_path = "#{Rails.root}/tmp/converted_#{SecureRandom.hex(8)}.html"
  File.write(html_path, html_content)
  html_path
end

2. 用Headless Selenium截图所有图表元素

这里解决了你之前直接打开docx的问题——先转成HTML,浏览器可以正常解析。定位.mammoth-drawing元素并截图,保存为Base64格式方便后续替换。

def capture_drawing_elements(html_path)
  # 配置Firefox Headless模式
  options = Selenium::WebDriver::Firefox::Options.new(binary: '/usr/bin/firefox', headless: true)
  driver = Selenium::WebDriver.for :firefox, options: options
  
  # 正确处理本地文件路径(转义特殊字符)
  encoded_path = ERB::Util.url_encode(html_path)
  driver.navigate.to("file://#{encoded_path}")
  
  # 定位所有drawing对应的HTML元素
  drawing_elements = driver.find_elements(:css, '.mammoth-drawing')
  screenshots = []
  
  drawing_elements.each_with_index do |element, index|
    img_path = "#{Rails.root}/tmp/drawing_#{index}.png"
    element.screenshot.save(img_path)
    # 转成Base64避免文件IO问题
    img_base64 = Base64.strict_encode64(File.read(img_path))
    screenshots << img_base64
    File.delete(img_path) # 清理临时文件
  end
  
  driver.quit
  screenshots
end

3. 将截图替换回原docx的drawing节点

解压docx包,修改word/document.xml,把所有<w:drawing>节点替换为标准的Word图片节点,同时更新文档的关系文件(确保图片能被正确引用)。

def replace_drawings_with_images(docx_path, screenshots)
  tmp_dir = "#{Rails.root}/tmp/docx_tmp_#{SecureRandom.hex(8)}"
  FileUtils.mkdir_p(tmp_dir)
  
  # 解压原docx到临时目录
  Zip::File.open(docx_path) do |zip|
    zip.each do |entry|
      entry.extract("#{tmp_dir}/#{entry.name}")
    end
  end
  
  # 读取并修改document.xml
  doc_xml_path = "#{tmp_dir}/word/document.xml"
  doc = Nokogiri::XML(File.read(doc_xml_path))
  drawing_nodes = doc.xpath('//w:drawing', 'w' => 'http://schemas.openxmlformats.org/wordprocessingml/2006/main')
  
  drawing_nodes.each_with_index do |node, index|
    next if index >= screenshots.size
    
    # 保存图片到docx的media目录
    img_filename = "media/image_#{index}.png"
    img_path = "#{tmp_dir}/word/#{img_filename}"
    File.write(img_path, Base64.strict_decode64(screenshots[index]))
    
    # 构建Word标准的图片XML节点,替换原drawing节点
    image_node = build_word_image_node(doc, img_filename, index)
    node.replace(image_node)
  end
  
  # 更新关系文件,添加图片引用
  update_rels_file("#{tmp_dir}/word/_rels/document.xml.rels", screenshots.size)
  
  # 重新打包成docx
  new_docx_path = "#{Rails.root}/tmp/modified_#{SecureRandom.hex(8)}.docx"
  Zip::File.open(new_docx_path, Zip::File::CREATE) do |zip|
    Dir.glob("#{tmp_dir}/**/*").each do |file|
      next if File.directory?(file)
      relative_path = file.sub("#{tmp_dir}/", '')
      zip.add(relative_path, file)
    end
  end
  
  # 清理临时目录
  FileUtils.rm_rf(tmp_dir)
  new_docx_path
end

# 辅助方法:构建符合Word OOXML标准的图片节点
def build_word_image_node(doc, img_filename, index)
  ns = {
    'w' => 'http://schemas.openxmlformats.org/wordprocessingml/2006/main',
    'r' => 'http://schemas.openxmlformats.org/officeDocument/2006/relationships',
    'wp' => 'http://schemas.openxmlformats.org/drawingml/2006/wordprocessingDrawing',
    'a' => 'http://schemas.openxmlformats.org/drawingml/2006/main',
    'pic' => 'http://schemas.openxmlformats.org/drawingml/2006/picture'
  }
  
  # 构建inline图片结构(Word中嵌入图片的标准格式)
  inline = doc.create_element('wp:inline', ns)
  inline.set_attributes({ distT: '0', distB: '0', distL: '0', distR: '0' })
  
  extent = doc.create_element('wp:extent', ns)
  extent.set_attributes({ cx: '914400', cy: '685800' }) # 1x0.75英寸,可根据截图调整
  inline.add_child(extent)
  
  doc_pr = doc.create_element('wp:docPr', ns)
  doc_pr.set_attributes({ id: "#{index + 1}", name: "Picture #{index + 1}" })
  inline.add_child(doc_pr)
  
  graphic = doc.create_element('a:graphic', ns)
  graphic_data = doc.create_element('a:graphicData', ns)
  graphic_data.set_attribute('uri', 'http://schemas.openxmlformats.org/drawingml/2006/picture')
  
  pic = doc.create_element('pic:pic', ns)
  nv_pic_pr = doc.create_element('pic:nvPicPr', ns)
  c_nv_pr = doc.create_element('pic:cNvPr', ns)
  c_nv_pr.set_attributes({ id: "#{index + 1}", name: "Picture #{index + 1}" })
  nv_pic_pr.add_child(c_nv_pr)
  pic.add_child(nv_pic_pr)
  
  blip_fill = doc.create_element('pic:blipFill', ns)
  blip = doc.create_element('a:blip', ns)
  blip.set_attribute('r:embed', "rId#{index + 10}") # rId要和rels文件对应
  blip_fill.add_child(blip)
  pic.add_child(blip_fill)
  
  graphic_data.add_child(pic)
  graphic.add_child(graphic_data)
  inline.add_child(graphic)
  
  # 包装在r节点中(Word段落中的元素容器)
  r = doc.create_element('w:r', ns)
  r.add_child(doc.create_element('w:rPr', ns))
  r.add_child(inline)
  
  r
end

# 辅助方法:更新rels文件,添加图片的关系引用
def update_rels_file(rels_path, count)
  doc = Nokogiri::XML(File.read(rels_path))
  ns = { 'r' => 'http://schemas.openxmlformats.org/package/2006/relationships' }
  
  # 获取现有最大rId,避免冲突
  existing_rids = doc.xpath('//r:Relationship/@Id', ns).map { |id| id.value.gsub('rId', '').to_i }
  max_rid = existing_rids.max || 0
  
  count.times do |i|
    rid = "rId#{max_rid + i + 1}"
    rel = doc.create_element('r:Relationship', ns)
    rel.set_attributes({
      Id: rid,
      Type: 'http://schemas.openxmlformats.org/officeDocument/2006/relationships/image',
      Target: "media/image_#{i}.png"
    })
    doc.root.add_child(rel)
  end
  
  File.write(rels_path, doc.to_xml)
end

4. 在上传流程中整合所有方法

# 假设这是你的上传处理方法,接收上传的文件对象
def process_uploaded_docx(uploaded_file)
  tmp_docx_path = "#{Rails.root}/tmp/uploaded_#{SecureRandom.hex(8)}.docx"
  File.write(tmp_docx_path, uploaded_file.read)
  
  begin
    html_path = convert_docx_to_html(tmp_docx_path)
    screenshots = capture_drawing_elements(html_path)
    modified_docx_path = replace_drawings_with_images(tmp_docx_path, screenshots)
    
    # 这里替换成你上传到Google Drive的逻辑
    # google_drive_service.upload_file(modified_docx_path, name: 'processed_document.docx')
    
    # 清理临时文件
    [tmp_docx_path, html_path, modified_docx_path].each do |path|
      File.delete(path) if File.exist?(path)
    end
  rescue => e
    Rails.logger.error("Docx processing failed: #{e.message}\n#{e.backtrace.join("\n")}")
    # 异常时清理临时文件
    [tmp_docx_path, html_path, modified_docx_path].each do |path|
      File.delete(path) if File.exist?(path)
    end
    raise e # 根据业务需求决定是否抛出异常
  end
end

关键优势

  1. 解决Selenium无法打开docx的问题:先转成HTML,浏览器可以正常解析渲染。
  2. 支持所有类型的drawing元素:不管是Excel图表、SVG还是复杂图形,只要能在HTML中渲染,就能被截图。
  3. 保留标准docx格式:替换后的文档符合Office Open XML规范,Google Drive可以正常处理,转HTML时会显示图片而非异常的drawing元素。

注意事项

  • 确保服务器上安装了Firefox和对应版本的Geckodriver(适配Selenium 3.142.7)。
  • 可以根据截图的实际尺寸调整图片节点的cx和cy属性(1英寸=914400 EMU单位)。
  • 处理大型文档时,建议增加Selenium的超时时间,避免截图失败。

内容的提问来源于stack exchange,提问作者altose87

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 10:05:26