You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Aspose.Words按Word显示顺序提取DOCX所有文本?

DOCX文本提取顺序混乱问题解决方案

问题根源:XML存储顺序≠显示顺序

DOCX的document.xml并不是按视觉显示顺序存储文本的,而是按元素类型分组存储:

  • 普通正文段落按文档流顺序存储,但浮动对象(文本框、框架、形状等)会被包装在<w:anchor>或<w:inline>节点中,这些节点可能被插入到<w:body>的任意位置(比如开头、关联段落附近),而非视觉对应的位置。
  • 你提到的relativeHeight属性只是控制浮动对象锚点与段落的垂直关联方式,和XML中的存储顺序无关。Word的显示顺序是由布局引擎动态计算的坐标决定的,和XML节点的存储顺序没有直接对应关系。

DOCX文本存储的核心规则

没有所谓的"优先级",只是分组逻辑:

  1. 普通 inline 内容(正文段落、行内图片)按文档流顺序存储。
  2. 浮动对象(文本框、框架、浮动图片)作为独立的锚点节点存储,位置可能和其视觉位置脱节,解析时会被优先读取,导致提取顺序混乱。

按视觉顺序提取文本的可行方案(Aspose.Words C++)

直接操作document.xml几乎不可能实现按视觉顺序提取,因为XML中没有存储布局坐标。必须借助Aspose.Words的布局API获取每个文本块的实际显示坐标,再排序输出。

关键思路

  1. 遍历文档中所有文本容器:包括普通段落、文本框、框架内的段落。
  2. 用LayoutCollector和LayoutEnumerator获取每个段落的左上角坐标(X,Y)(通过RectangleF属性)。
  3. 按「Y坐标升序(从上到下)→ X坐标升序(从左到右)」的规则排序所有文本块。
  4. 按排序后的顺序输出文本。

代码实现示例

#include <Aspose.Words.Cpp/Document.h>
#include <Aspose.Words.Cpp/Layout/LayoutCollector.h>
#include <Aspose.Words.Cpp/Layout/LayoutEnumerator.h>
#include <Aspose.Words.Cpp/Layout/LayoutEntityType.h>
#include <Aspose.Words.Cpp/Nodes/NodeType.h>
#include <Aspose.Words.Cpp/Nodes/CompositeNode.h>
#include <Aspose.Words.Cpp/Shapes/Shape.h>
#include <Aspose.Words.Cpp/Frames/Frame.h>
#include <vector>
#include <algorithm>

using namespace Aspose::Words;
using namespace Aspose::Words::Layout;
using namespace Aspose::Words::Shapes;
using namespace Aspose::Words::Frames;
using namespace System::Drawing;

// 存储文本块及其布局坐标的结构体
struct TextBlockWithPosition
{
    System::String Text;
    RectangleF Bounds;
};

// 遍历所有节点,收集所有文本块(包括浮动对象内的)
void CollectAllTextBlocks(SharedPtr<Node> node, SharedPtr<LayoutCollector> collector, std::vector<TextBlockWithPosition>& textBlocks)
{
    if (node->get_NodeType() == NodeType::Paragraph)
    {
        auto para = System::DynamicCast<Paragraph>(node);
        if (para->get_Range()->get_Text().Trim().Length() == 0)
            return; // 跳过空段落

        // 获取段落的布局边界
        auto enumerator = MakeObject<LayoutEnumerator>(para->get_Document());
        enumerator->Current = collector->GetEntity(para);
        RectangleF bounds = enumerator->get_Rectangle();

        textBlocks.push_back({ para->get_Range()->get_Text(), bounds });
        return;
    }
    else if (node->get_NodeType() == NodeType::Shape)
    {
        auto shape = System::DynamicCast<Shape>(node);
        if (shape->get_HasText())
        {
            // 遍历文本框内的段落
            for (auto para : shape->get_TextFrame()->get_Paragraphs())
                CollectAllTextBlocks(para, collector, textBlocks);
        }
    }
    else if (node->get_NodeType() == NodeType::Frame)
    {
        auto frame = System::DynamicCast<Frame>(node);
        // 遍历框架内的段落
        for (auto para : frame->get_Paragraphs())
            CollectAllTextBlocks(para, collector, textBlocks);
    }

    // 递归遍历子节点
    if (System::DynamicCast<CompositeNode>(node) != nullptr)
    {
        auto compositeNode = System::DynamicCast<CompositeNode>(node);
        for (auto child : compositeNode->get_ChildNodes())
            CollectAllTextBlocks(child, collector, textBlocks);
    }
}

int main()
{
    // 加载文档
    auto doc = MakeObject<Document>(u"input.docx");
    
    // 初始化布局收集器
    auto collector = MakeObject<LayoutCollector>(doc);
    
    // 收集所有文本块
    std::vector<TextBlockWithPosition> textBlocks;
    CollectAllTextBlocks(doc->get_ChildNodes()->GetEnumerator()->get_Current(), collector, textBlocks);
    
    // 按视觉顺序排序:先按Y坐标(从上到下),再按X坐标(从左到右)
    std::sort(textBlocks.begin(), textBlocks.end(), [](const TextBlockWithPosition& a, const TextBlockWithPosition& b) {
        if (a.Bounds.Y != b.Bounds.Y)
            return a.Bounds.Y < b.Bounds.Y;
        return a.Bounds.X < b.Bounds.X;
    });
    
    // 输出到文件
    System::IO::StreamWriter writer(u"output.txt");
    for (auto& block : textBlocks)
        writer.WriteLine(block.Text.Trim());
    writer.Close();
    
    return 0;
}

关键说明

  • 弃用的get_X()/get_Y()可以用LayoutEnumerator::get_Rectangle()替代,返回的RectangleF包含了左上角的X、Y坐标和宽高。
  • 确保遍历所有节点类型:Paragraph、Shape(文本框)、Frame,避免遗漏文本。
  • 排序逻辑中,Y坐标越小越靠上(Word的坐标系原点在页面左上角),所以按Y升序排列就是从上到下;Y相同时按X升序就是从左到右。

为什么直接操作XML不可行?

document.xml中没有存储任何布局坐标信息,所有视觉位置都是Word布局引擎根据文档内容、样式、页面设置动态计算的。只有通过Aspose.Words这类封装了布局引擎的API,才能获取到准确的显示坐标,进而实现按视觉顺序提取。

内容的提问来源于stack exchange,提问作者vignesh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 04:43:17