You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python脚本仅获取前两份文档,如何获取产品说明书与保修详情?

问题:筛选JSON API返回的指定文档并修改Python脚本

我是Python新手,当前脚本从JSON API调用中只能获取前两份文档,但业务需要优先获取Product Sheet(产品说明书)和Warranty Details(保修详情),同时能按需获取其他文档,求修改脚本的方法。

JSON片段示例

"documents":{
  "en":[
    {"label":"Installation guide (en)","type":"Installation guide","url":"https://media.datatail.com/docs/installation/401298_en.pdf"},
    {"label":"Owner's manual (en)","type":"Owner's manual","url":"https://media.datatail.com/docs/manual/401298_en.pdf"},
    {"label":"Product Sheet (en)","type":"Product specification","url":"https://media.datatail.com/docs/specs/401298_en.pdf"},
    {"label":"Warranty Details (en)","type":"Warranty Details","url":"https://media.datatail.com/docs/warranty/401298_en.pdf"},
    {"label":"Dimensions Guide (en)","type":"Dimension sheet","url":"https://media.datatail.com/docs/dimensions/401298_en.pdf"}
  ]
}

JSON配置请求

],
    "Body HTML": [
      "-",
      "description.en",
      "features.en.0.title",
      "document.en.0.label"
    ]

Python请求代码

# Columns with non-standard mapping or formatting are populated below. They are:
# 1. "Body HTML"
# 2. "Metafield: sf_classifications.product_details [string]"
# 3. "Image Src"

# 1. "Body HTML"
# Multiple and diverse values written in HTML format

for item in data['data']:
    specifications_en = item.get('specifications', {}).get('en', [])
    for section in specifications_en:
        section_title = section.get('section', '')

        if section_title == "Features":
            values = section.get('values', [])

            if values:
                features = '<p><b>Features</b></p><ul>'
                for spec in values:
                    specification = spec.get('specification', '')
                    value = spec.get('value', '')
                    unit = spec.get('unit', '')
                    features += f"\n<li>{specification}:<br/>{value} {unit}</li>"

                features += '</ul>'

for item in data['data']:
    documents_en = item.get('documents', {}).get('en', [])

    if documents_en:
        documents = '<p><b>Documents</b></p><ul>'
        for document in documents_en:
            label = document.get('label', '')
            url = document.get('url', '').replace('\/', '/')
            documents += f"\n<li><a href='{url}'>{label}</a></li>"

        documents += '</ul>'

row[header_indices["Body HTML"]] = '<body>\n<p>'
row[header_indices["Body HTML"]] += obj.get('description').get('en')
row[header_indices["Body HTML"]] += '</p>\n'
row[header_indices["Body HTML"]] += features
row[header_indices["Body HTML"]] += documents
row[header_indices["Body HTML"]] += '\n</body>'

修改方案

核心逻辑

在遍历文档列表时,通过标签筛选出需要的文档,可优先保留目标文档,也可按需补充其他文档。

修改后的文档处理代码

# 定义需要获取的文档标签列表,按需添加
required_docs = ["Product Sheet (en)", "Warranty Details (en)"]

for item in data['data']:
    documents_en = item.get('documents', {}).get('en', [])

    if documents_en:
        documents = '<p><b>Documents</b></p><ul>'
        # 筛选出目标文档
        filtered_docs = [doc for doc in documents_en if doc.get('label') in required_docs]
        
        # 可选:如果需要同时保留其他文档,取消以下两行注释
        # other_docs = [doc for doc in documents_en if doc.get('label') not in required_docs]
        # filtered_docs.extend(other_docs)
        
        for document in filtered_docs:
            label = document.get('label', '')
            url = document.get('url', '').replace('\/', '/')
            documents += f"\n<li><a href='{url}'>{label}</a></li>"

        documents += '</ul>'

说明

  1. 把需要的文档标签写入required_docs列表,脚本会自动筛选出对应文档
  2. 若需保留其他文档,取消注释other_docs和extend代码,目标文档会优先显示
  3. 原脚本未做筛选,所以会输出所有文档,通过列表推导式即可实现精准筛选

内容的提问来源于stack exchange,提问作者Irene C

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.30 12:47:12