You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python检测Docx失效链接脚本报错:'list'对象无'items'属性

问题排查与修复

错误原因

find_broken_links_in_docx函数返回的是失效链接的列表,但write_report函数预期接收的是键为文件路径、值为对应失效链接列表的字典,调用items()遍历列表自然会触发AttributeError。

此外脚本还有其他潜在问题:

  • 未导入必需的requests模块(代码里用到了requests.head)
  • 导入了无用的url模块
  • find_hyperlinks会收集所有带"hyperlink"的关联,包括文档内部链接,这些链接不需要检测有效性
  • 输出目录不存在时会写入失败

修复后的完整代码

import docx
import os
import requests

def find_hyperlinks(doc):
    hyperlinks = []
    rels = doc.part.rels
    for rel_id in rels:
        rel = rels[rel_id]
        # 只收集外部HTTP/HTTPS链接,排除内部文档链接
        if rel.target_ref.startswith(('http://', 'https://')):
            hyperlinks.append(rel.target_ref)
    return hyperlinks

def find_broken_links_in_docx(file_path):
    doc = docx.Document(file_path)
    broken_links = []
    hyperlinks = find_hyperlinks(doc)
    for link in hyperlinks:
        try:
            # 设置超时避免长时间等待
            response = requests.head(link, allow_redirects=True, timeout=5)
            if response.status_code >= 400:
                broken_links.append(link)
        except requests.RequestException:
            broken_links.append(link)
    # 返回字典格式,适配write_report的预期
    return {file_path: broken_links}

def write_report(report, output_file):
    # 确保输出目录存在
    output_dir = os.path.dirname(output_file)
    if not os.path.exists(output_dir):
        os.makedirs(output_dir)
    
    with open(output_file, 'w', encoding='utf-8') as f:
        for file_path, links in report.items():
            f.write(f"File: {file_path}\n")
            if not links:
                f.write("  No broken links found.\n")
            else:
                for link in links:
                    f.write(f"  Broken link: {link}\n")
            f.write("\n")

if __name__ == "__main__":
    target_doc = "C:/Users/demo.docx"
    output_file = "C:/Results/broken_links_report.txt"
    report = find_broken_links_in_docx(target_doc)
    write_report(report, output_file)
    print(f"Report written to {output_file}")

关键改动说明

  1. 修正数据结构匹配:
    • find_broken_links_in_docx现在接收文件路径参数,返回{文件路径: 失效链接列表}的字典,完美适配write_report的遍历逻辑。
  2. 补全依赖与清理冗余:
    • 添加import requests,移除无用的url模块和allText变量。
  3. 优化链接筛选逻辑:
    • find_hyperlinks只保留以http://或https://开头的外部链接,避免检测文档内部的锚点链接。
  4. 增加鲁棒性:
    • 给requests.head添加timeout=5,防止因链接超时导致脚本挂起。
    • 自动创建输出目录,避免因目录不存在导致的写入失败。
    • 报告中增加"无失效链接"的提示,输出更友好。
  5. 代码结构优化:
    • 将文档初始化移到find_broken_links_in_docx内部,让函数更独立通用,方便后续扩展为批量处理多个文档。

内容的提问来源于stack exchange,提问作者Travis Webb

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.14 00:03:10