You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

C#实现Word转PDF时部分内部文档链接未被修改的排查求助

C#实现Word转PDF时部分内部文档链接未被修改的排查求助

我最近在做一个C#工具,功能是把Word(.doc格式)转成PDF,同时要递归处理所有关联的文档——不仅要把这些关联文档也转成PDF,还要把原文档里所有指向这些.doc/.docx/.dotx的链接都替换成对应的PDF路径。

我写了一段处理链接的代码,但遇到了棘手的问题:总有36个链接没被修改,我翻遍了能想到的地方都找不到它们藏在哪。先贴一下我的处理代码:

bool continueing = true;
toc = WordprocessingDocument.Open(inFile, true);
HashSet<string> seen_hrs = new HashSet<string>();
while (continueing)
{
    continueing = false;
    IEnumerable<HyperlinkRelationship> all_hr = toc.MainDocumentPart.HyperlinkRelationships;
    foreach (HyperlinkRelationship hr in all_hr)
    {
        if (seen_hrs.Contains(hr.Id)) continue;
        seen_hrs.Add(hr.Id);
        string link = hr.Uri.OriginalString.Replace("%20", " ");
        link = link.Trim();
        string link_originalName = link;
        if (! (link.EndsWith(".doc") || link.EndsWith(".docx") || link.EndsWith(".dotx"))) continue;
        if (link.Contains("http")) continue;
        string filename = Path.GetFileName(hr.Uri.OriginalString);
        filename = filename.Replace("%20", " ");
        filename = filename.Trim();
        link = GetLink(filename);
        if (link == "") continue;
        link_originalName = link;
        var hyperlinkRelationshipId = hr.Id;
        var pos = link.LastIndexOf('/');
        if (pos == -1) pos = link.LastIndexOf("\\");
        var pos_point = filename.LastIndexOf('.');
        var filestem = filename.Substring(0, pos_point);
        string base_path = link.Substring(0, link.Length - Path.GetFileName(link_originalName).Length);
        filestem = filestem.Replace('.', '_');
        link = base_path + filestem + ".pdf";
        toc.MainDocumentPart.DeleteReferenceRelationship(hr);
        try
        {
            toc.MainDocumentPart.AddHyperlinkRelationship(new Uri(link, UriKind.Relative), false, hyperlinkRelationshipId);
            dependents.Add(Tuple.Create(path + "\\" + link_originalName, outputPath + "\\" + link));
        }
        catch (Exception)
        {
            ui.AddError("Wrong link: " + hr.Uri.OriginalString + " in file: " + pathFile + " that could not be corrected.");
        }
        continueing = true;
        break;
    }
}

目前我已经做了这些排查:

  • toc.MainDocumentPart.HyperlinkRelationships里的439个链接都被成功处理了,但还有36个链接没被捕获到
  • 检查了toc.MainDocumentPart.ExternalRelationships.Count(),结果是0
  • toc.HyperlinkRelationships.Count()也是0,toc.RootPart.HyperlinkRelationships的数量和主文档部分一致,还是439
  • 文档里没有形状(Shapes)相关的元素,所有链接用Alt+F9显示后看起来都是普通的文本链接,没有特殊格式

实在找不到这36个漏网链接的藏身之处了,有没有大佬遇到过类似情况?这些未被处理的链接可能藏在Word文档结构的哪些地方呢?

内容来源于stack exchange

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.08 07:29:32