You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提取PDF表格并转换为支持rowspan/colspan的HTML表格?

Hey, the problem with your current code is that extracting raw text discards all the table's structural information—things like merged cells (rowspan/colspan) and column/row alignment. To generate the HTML table you need, we need to use methods that can detect and preserve this layout. Below are two practical approaches for C#:

Solution: Convert PDF Tables to HTML with Rowspan/Colspan

TabulaSharp is a .NET library built specifically for extracting tables from PDFs while preserving their structure, including rowspan and colspan attributes. It’s the easiest way to handle this without writing complex layout-parsing logic.

Step 1: Install TabulaSharp

First, add the NuGet package to your project:

Install-Package TabulaSharp

Step 2: Extract Tables and Convert to HTML

Here’s code to extract tables from your PDF, convert them to your desired HTML format, and save the result to a TXT file:

using System;
using System.IO;
using System.Text;
using System.Collections.Generic;
using Tabula;
using Tabula.Utilities;

class Program
{
    static void Main(string[] args)
    {
        string pdfPath = @"C:\Temp\Y.pdf";
        string outputPath = @"C:\Temp\Y.txt";
        
        // Extract structured tables from the PDF
        var tables = ExtractTablesFromPdf(pdfPath);
        
        // Convert tables to the target HTML format
        string htmlOutput = ConvertTablesToHtml(tables);
        
        // Save to TXT file
        File.WriteAllText(outputPath, htmlOutput);
        
        Console.WriteLine("Done!");
        Console.ReadLine();
    }
    
    private static List<Table> ExtractTablesFromPdf(string pdfPath)
    {
        var tables = new List<Table>();
        using (var stream = new FileStream(pdfPath, FileMode.Open, FileAccess.Read))
        {
            var reader = new PdfReader(stream);
            for (int pageNum = 1; pageNum <= reader.NumberOfPages; pageNum++)
            {
                // Extract all tables from the current page
                var pageTables = TableExtractor.Extract(reader, pageNum);
                tables.AddRange(pageTables);
            }
        }
        return tables;
    }
    
    private static string ConvertTablesToHtml(List<Table> tables)
    {
        var htmlBuilder = new StringBuilder();
        
        foreach (var table in tables)
        {
            htmlBuilder.AppendLine("<table frame=\"box\">");
            htmlBuilder.AppendLine("  <colgroup>");
            
            // Calculate total columns (accounting for colspans)
            int totalColumns = CalculateTotalColumns(table);
            htmlBuilder.AppendLine($"    <col span=\"{totalColumns}\"/>");
            
            htmlBuilder.AppendLine("  </colgroup>");
            htmlBuilder.AppendLine("  <tbody>");
            
            foreach (var row in table.Rows)
            {
                htmlBuilder.AppendLine("    <tr>");
                
                foreach (var cell in row.Cells)
                {
                    var cellAttributes = new StringBuilder();
                    if (cell.RowSpan > 1)
                        cellAttributes.Append($" rowspan=\"{cell.RowSpan}\"");
                    if (cell.ColSpan > 1)
                        cellAttributes.Append($" colspan=\"{cell.ColSpan}\"");
                    
                    // Clean up cell text (remove newlines and extra spaces)
                    string cleanText = cell.Text.Replace("\n", " ").Trim();
                    htmlBuilder.AppendLine($"      <td{cellAttributes}>{cleanText}</td>");
                }
                
                htmlBuilder.AppendLine("    </tr>");
            }
            
            htmlBuilder.AppendLine("  </tbody>");
            htmlBuilder.AppendLine("</table>");
        }
        
        return htmlBuilder.ToString();
    }
    
    private static int CalculateTotalColumns(Table table)
    {
        int maxColumns = 0;
        foreach (var row in table.Rows)
        {
            int rowColumnCount = 0;
            foreach (var cell in row.Cells)
                rowColumnCount += cell.ColSpan;
            if (rowColumnCount > maxColumns)
                maxColumns = rowColumnCount;
        }
        return maxColumns;
    }
}

Approach 2: Manual Parsing with iTextSharp (Advanced)

If you want to stick with iTextSharp, you’ll need to analyze the positional data of text elements to reconstruct the table structure. This requires more work but gives you full control over the parsing logic.

Step 1: Extract Text with Position Data

Modify your existing code to capture each text element’s coordinates (critical for determining table layout):

using System;
using System.Collections.Generic;
using iTextSharp.text.pdf;
using iTextSharp.text.pdf.parser;

public class TextWithPosition
{
    public string Text { get; set; }
    public float X { get; set; }
    public float Y { get; set; }
    public float Width { get; set; }
}

public static List<TextWithPosition> GetTextWithPositionsFromPDF(string path)
{
    var textElements = new List<TextWithPosition>();
    
    using (PdfReader reader = new PdfReader(path))
    {
        for (int pageNum = 1; pageNum <= reader.NumberOfPages; pageNum++)
        {
            var listener = new PositionTextExtractionStrategy();
            PdfTextExtractor.GetTextFromPage(reader, pageNum, listener);
            
            foreach (var renderInfo in listener.GetRenderInfos())
            {
                string text = renderInfo.GetText();
                if (!string.IsNullOrWhiteSpace(text))
                {
                    var rect = renderInfo.GetBaseline().GetBoundingRectange();
                    textElements.Add(new TextWithPosition
                    {
                        Text = text,
                        X = rect.Left,
                        Y = rect.Bottom,
                        Width = rect.Width
                    });
                }
            }
        }
    }
    
    return textElements;
}

Step 2: Reconstruct Table Structure

To build the table:

  1. Sort text elements by descending Y (PDF coordinates start at the bottom-left corner).
  2. Group elements into rows by their Y position (use a small tolerance to account for minor alignment differences).
  3. For each row, sort elements by X position to get columns.
  4. Detect merged cells:
    • Rowspan: Cells with the same X range that appear across multiple rows.
    • Colspan: Adjacent cells in the same row with no gaps between them (indicating a merged column).

Step 3: Generate HTML Table

Once you have the structured table data (with rowspan/colspan metadata), you can generate the HTML using the same logic as the TabulaSharp approach.

Key Notes

  • TabulaSharp is the better choice for most cases—it handles edge cases like irregular columns and merged cells out of the box.
  • If your PDF has extremely complex tables (e.g., nested tables), you may need to adjust the code to handle those scenarios.
  • Always test with your specific PDF, as table extraction can vary based on how the PDF was generated.

内容的提问来源于stack exchange,提问作者CD DelRio

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 07:44:46