如何使用Prawn Gem在PDF中正确渲染HTML标签
问题
在Rails应用中使用prawn gem生成PDF报告时,无法正确渲染包含<b>、<ul>、<li>、<img>、<table>等HTML标签的描述数据,这些标签会直接以文本形式显示在PDF中,而非解析成对应的格式。相关代码如下:
# frozen_string_literal: true require 'prawn' require 'prawn/table' module User module Reports class GeneratePdf include ActionView::Helpers::NumberHelper include ActionView::Helpers::DateHelper include Seo::ApplicationHelper BORDER_COLOR = 'dddddd'.freeze TABLE_HEADER_COLOR = 'f2f0f0'.freeze WHITE_COLOR = 'ffffff'.freeze TABLE_HEADERS = [[I18n.t("seo.reports.type_of_fix"), I18n.t("seo.reports.description"), I18n.t("seo.reports.fixed_at")]].freeze A4_SIZE = [595.276, 841.890] attr_reader :shop def initialize(shop) @shop = shop @pdf = Prawn::Document.new(page_size: A4_SIZE) @document_width = @pdf.bounds.width end def perform report_summary_section reports_table @pdf.render end private def formatted_reports_data reports = Report.all reports.map do |report| fix_type = if report.issue_type.include?('_created') pretty_report_issue_type(report.issue_type, report.get_description_count) else pretty_report_issue_type(report.issue_type) end description = pretty_report_description(report.issue_type, report.description) # 此处为带HTML标签的内容 fixed_at = "#{distance_of_time_in_words_to_now(report.fixed_at)} ago" if report.fixed_at [fix_type, description, fixed_at] end end def report_summary_section @pdf.text I18n.t("seo.dashboard.reports.title"), size: 16, style: :bold @pdf.move_down 5 @pdf.text @shop.domain, size: 12, color: 'a6a6a6', inline_format: true @pdf.move_down 10 summary_data = [ [ "#{I18n.t("seo.dashboard.reports.total_fixes")}\n<b>#{number_with_delimiter(@shop.reports_count) || 0}</b>", "#{I18n.t("seo.dashboard.reports.title_tags_fixed")}\n<b>#{number_with_delimiter(@shop.report_title_count) || 0}</b>", "#{I18n.t("seo.dashboard.reports.meta_tags_fixed")}\n<b>#{number_with_delimiter(@shop.report_desc_count) || 0}</b>" ], [ "#{I18n.t("seo.dashboard.reports.alt_texts_fixed")}\n<b>#{number_with_delimiter(@shop.report_alt_count) || 0}</b>", "#{I18n.t("seo.dashboard.reports.links_fixed")}\n<b>#{number_with_delimiter(@shop.report_links_count) || 0}</b>", "#{I18n.t("seo.dashboard.reports.bulk_updates")}\n<b>#{number_with_delimiter(@shop.report_bulk_count) || 0}</b>" ] ] column_width = @document_width / 3.0 summary_options = { cell_style: { size: 12, inline_format: true, padding: [15, 10], borders: [:top, :bottom, :left, :right], border_width: 1, border_color: BORDER_COLOR, align: :center, valign: :center }, column_widths: [column_width] * 3 } @pdf.table(summary_data, summary_options) do |table| table.cells.each do |cell| cell.content = cell.content.sub('<b>', '<font size="18"><b>').sub('</b>', '</b></font>') end end end def reports_table @pdf.move_down 20 report_data_options = { width: @document_width, row_colors: [WHITE_COLOR], cell_style: { border_width: 1, borders: [:bottom], border_color: BORDER_COLOR, size: 10, inline_format: true } } @pdf.table(TABLE_HEADERS + formatted_reports_data, report_data_options) do |table| table.row(0).font_style = :bold table.row(0).background_color = TABLE_HEADER_COLOR table.row(0).size = 9 table.cells.each do |cell| cell.content = cell.content # 计划在此处处理HTML标签,但未找到可行方案 end end end end end end
解决方案
方法1:使用prawn-html gem扩展HTML解析能力
Prawn原生仅支持少量inline格式标签,prawn-html是专门为其开发的HTML渲染插件,支持<b>、<ul>、<li>、<img>、<table>等常见标签。
- 添加gem依赖:
# Gemfile gem 'prawn-html'
执行bundle install安装。
- 修改表格单元格渲染逻辑:
在reports_table方法中,针对描述列(第二列,索引为1)替换原有单元格内容的渲染方式:
table.cells.each do |cell| # 仅处理描述列 next unless cell.column == 1 # 清空默认文本内容,改用HTML渲染 cell.content = '' @pdf.bounding_box(cell.bounds, width: cell.width, height: cell.height) do @pdf.html(cell.content_was, inline_format: true) end end
此方法会自动解析HTML标签并转换成PDF对应的格式。
方法2:手动解析HTML标签(自定义场景)
如果不想引入新gem,可以针对需要的标签手动解析转换:
- 添加HTML处理辅助方法:
private def render_html_content(pdf, html) # 处理粗体标签 html = html.gsub(/<b>(.*?)<\/b>/, '<b>\1</b>') # 处理无序列表 html.split(/<ul>|<\/ul>/).each do |section| next if section.strip.empty? if section.include?('<li>') section.split(/<li>|<\/li>/).reject(&:empty?).each do |item| pdf.text "• #{item}", inline_format: true end else pdf.text section, inline_format: true end end # 扩展处理图片:解析<img>标签,提取src路径后用pdf.image渲染 # 扩展处理表格:解析<table>结构,用Prawn::Table重新构建 end
- 在表格中调用该方法:
table.cells.each do |cell| next unless cell.column == 1 cell.content = '' @pdf.bounding_box(cell.bounds, width: cell.width, height: cell.height) do render_html_content(@pdf, cell.content_was) end end
此方法需要根据实际需求扩展标签处理逻辑,适合高度自定义的场景。
内容的提问来源于stack exchange,提问作者aldrien.h
相关产品推荐
相关产品推荐

