You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用C#和OpenXML SDK无内存加载处理大DocX表格数据添加

Handling Large DOCX Files with Streaming (SAX-Style) OpenXML SDK

Hey, this is a super common pain point when working with large DOCX files—your current DOM-based approach loads the entire document tree into memory, which is why you’re hitting memory limits with 100MB+ files. The good news is you absolutely can use a SAX-style streaming approach with the OpenXML SDK to avoid loading the full DOM.

Why Your Current Approach Fails

The WordprocessingDocument.Open + DOM traversal method pulls every element of the document into memory at once. For small files this is fine, but large documents (especially those with big tables) will eat up RAM quickly.

The Streaming Solution: OpenXmlReader + OpenXmlWriter

The OpenXML SDK includes OpenXmlReader and OpenXmlWriter specifically for this scenario. These classes let you read/write the document element-by-element, only keeping the current element in memory. Here’s how to adapt your code to use this approach:

Step-by-Step Code Implementation

using System.IO;
using DocumentFormat.OpenXml;
using DocumentFormat.OpenXml.Packaging;
using DocumentFormat.OpenXml.Wordprocessing;

public void AddRowsToLargeDocxTable(string sourceFilePath, string targetFilePath, IEnumerable<IEnumerable<string>> data)
{
    // Open source document in read-only mode
    using (var sourceDoc = WordprocessingDocument.Open(sourceFilePath, false))
    // Create target document (can be a temp file)
    using (var targetDoc = WordprocessingDocument.Create(targetFilePath, WordprocessingDocumentType.Document))
    {
        // Copy critical parts from source to target (styles, fonts, settings)
        targetDoc.AddMainDocumentPart();
        sourceDoc.MainDocumentPart.DocumentSettingsPart?.CopyTo(targetDoc.MainDocumentPart.AddNewPart<DocumentSettingsPart>());
        sourceDoc.MainDocumentPart.StyleDefinitionsPart?.CopyTo(targetDoc.MainDocumentPart.AddNewPart<StyleDefinitionsPart>());
        sourceDoc.MainDocumentPart.FontTablePart?.CopyTo(targetDoc.MainDocumentPart.AddNewPart<FontTablePart>());

        var sourceMainPart = sourceDoc.MainDocumentPart;
        var targetMainPart = targetDoc.MainDocumentPart;

        using (var reader = OpenXmlReader.Create(sourceMainPart))
        using (var writer = OpenXmlWriter.Create(targetMainPart))
        {
            bool isTargetTableProcessed = false;

            while (reader.Read())
            {
                // Handle root Document element
                if (reader.IsStartElement && reader.ElementType == typeof(Document))
                {
                    writer.WriteStartElement(reader);
                    continue;
                }

                // Handle Body element
                if (reader.IsStartElement && reader.ElementType == typeof(Body))
                {
                    writer.WriteStartElement(reader);
                    continue;
                }

                // Target the first Table (matches your original .First() logic)
                if (reader.IsStartElement && reader.ElementType == typeof(Table) && !isTargetTableProcessed)
                {
                    isTargetTableProcessed = true;
                    // Write the Table's start tag
                    writer.WriteStartElement(reader);

                    // Read and write all existing content inside the Table
                    while (reader.Read())
                    {
                        // Insert new rows right before the Table's end tag
                        if (reader.IsEndElement && reader.ElementType == typeof(Table))
                        {
                            // Add your custom rows (reusing your original row creation logic)
                            foreach (var row in data)
                            {
                                var tableRow = new TableRow();
                                foreach (var cellValue in row)
                                {
                                    var tableCell = new TableCell(
                                        new Paragraph(
                                            new Run(
                                                new Text(cellValue)
                                            )
                                        )
                                    );
                                    // Optional: Copy cell properties from existing rows to match formatting
                                    tableRow.Append(tableCell);
                                }
                                writer.WriteElement(tableRow);
                            }
                            // Write the Table's end tag
                            writer.WriteEndElement();
                            break;
                        }

                        // Pass through existing Table child elements
                        if (reader.IsStartElement)
                        {
                            writer.WriteStartElement(reader);
                        }
                        else if (reader.IsEndElement)
                        {
                            writer.WriteEndElement();
                        }
                        else if (reader.IsText)
                        {
                            writer.WriteString(reader.Value);
                        }
                    }
                }
                else
                {
                    // Copy all other document elements as-is
                    if (reader.IsStartElement)
                    {
                        writer.WriteStartElement(reader);
                    }
                    else if (reader.IsEndElement)
                    {
                        writer.WriteEndElement();
                    }
                    else if (reader.IsText)
                    {
                        writer.WriteString(reader.Value);
                    }
                }
            }
        }
    }

    // Optional: Replace original file with modified version (always back up first!)
    File.Delete(sourceFilePath);
    File.Move(targetFilePath, sourceFilePath);
}

Key Notes & Improvements

  • Memory Efficiency: This code only holds the current element in memory, so it works seamlessly with even gigabyte-sized DOCX files.
  • Table Targeting: The example targets the first table (like your original code). If you need a specific table, add logic to check for table IDs, styles, or content before processing.
  • Formatting: If your existing table has cell widths, styling, or formatting, copy the TableCellProperties from an existing row to your new rows to match the document’s look.
  • Safety: Always back up your original file before replacing it! Add try/catch blocks to handle file locks or unexpected document structures.

When to Use This Approach

Use streaming whenever:

  • Your DOCX file is larger than 50MB (DOM starts to struggle here)
  • You’re modifying specific parts of a document without needing access to the entire tree
  • Memory usage is a critical constraint (e.g., server environments with limited RAM)

内容的提问来源于stack exchange,提问作者Tom Donkers

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 09:17:32