Python脚本仅获取前两份文档,如何获取产品说明书与保修详情?
问题:筛选JSON API返回的指定文档并修改Python脚本
我是Python新手,当前脚本从JSON API调用中只能获取前两份文档,但业务需要优先获取Product Sheet(产品说明书)和Warranty Details(保修详情),同时能按需获取其他文档,求修改脚本的方法。
JSON片段示例
"documents":{ "en":[ {"label":"Installation guide (en)","type":"Installation guide","url":"https://media.datatail.com/docs/installation/401298_en.pdf"}, {"label":"Owner's manual (en)","type":"Owner's manual","url":"https://media.datatail.com/docs/manual/401298_en.pdf"}, {"label":"Product Sheet (en)","type":"Product specification","url":"https://media.datatail.com/docs/specs/401298_en.pdf"}, {"label":"Warranty Details (en)","type":"Warranty Details","url":"https://media.datatail.com/docs/warranty/401298_en.pdf"}, {"label":"Dimensions Guide (en)","type":"Dimension sheet","url":"https://media.datatail.com/docs/dimensions/401298_en.pdf"} ] }
JSON配置请求
], "Body HTML": [ "-", "description.en", "features.en.0.title", "document.en.0.label" ]
Python请求代码
# Columns with non-standard mapping or formatting are populated below. They are: # 1. "Body HTML" # 2. "Metafield: sf_classifications.product_details [string]" # 3. "Image Src" # 1. "Body HTML" # Multiple and diverse values written in HTML format for item in data['data']: specifications_en = item.get('specifications', {}).get('en', []) for section in specifications_en: section_title = section.get('section', '') if section_title == "Features": values = section.get('values', []) if values: features = '<p><b>Features</b></p><ul>' for spec in values: specification = spec.get('specification', '') value = spec.get('value', '') unit = spec.get('unit', '') features += f"\n<li>{specification}:<br/>{value} {unit}</li>" features += '</ul>' for item in data['data']: documents_en = item.get('documents', {}).get('en', []) if documents_en: documents = '<p><b>Documents</b></p><ul>' for document in documents_en: label = document.get('label', '') url = document.get('url', '').replace('\/', '/') documents += f"\n<li><a href='{url}'>{label}</a></li>" documents += '</ul>' row[header_indices["Body HTML"]] = '<body>\n<p>' row[header_indices["Body HTML"]] += obj.get('description').get('en') row[header_indices["Body HTML"]] += '</p>\n' row[header_indices["Body HTML"]] += features row[header_indices["Body HTML"]] += documents row[header_indices["Body HTML"]] += '\n</body>'
修改方案
核心逻辑
在遍历文档列表时,通过标签筛选出需要的文档,可优先保留目标文档,也可按需补充其他文档。
修改后的文档处理代码
# 定义需要获取的文档标签列表,按需添加 required_docs = ["Product Sheet (en)", "Warranty Details (en)"] for item in data['data']: documents_en = item.get('documents', {}).get('en', []) if documents_en: documents = '<p><b>Documents</b></p><ul>' # 筛选出目标文档 filtered_docs = [doc for doc in documents_en if doc.get('label') in required_docs] # 可选:如果需要同时保留其他文档,取消以下两行注释 # other_docs = [doc for doc in documents_en if doc.get('label') not in required_docs] # filtered_docs.extend(other_docs) for document in filtered_docs: label = document.get('label', '') url = document.get('url', '').replace('\/', '/') documents += f"\n<li><a href='{url}'>{label}</a></li>" documents += '</ul>'
说明
- 把需要的文档标签写入
required_docs列表,脚本会自动筛选出对应文档 - 若需保留其他文档,取消注释
other_docs和extend代码,目标文档会优先显示 - 原脚本未做筛选,所以会输出所有文档,通过列表推导式即可实现精准筛选
内容的提问来源于stack exchange,提问作者Irene C
相关产品推荐
相关产品推荐

