如何使用Python访问JSON-LD/Microdata列表提取产品元数据
实现方法
不要硬编码JSON-LD列表的下标位置(比如示例里写死取[3]的写法鲁棒性很差,页面结构调整后下标变了就会报错),正确逻辑是遍历json-ld数组下的所有条目,逐一判断条目的@type属性是否为Product,匹配到目标条目后再提取需要的字段即可。
可直接运行的代码
data = {'json-ld': [{'@context': 'http://schema.org', '@type': 'WebSite', 'potentialAction': {'@type': 'SearchAction', 'query-input': 'required ' 'name=search_term_string', 'target': 'https://www.vitenda.de/search/result?term={search_term_string}'}, 'url': 'https://www.vitenda.de'}, {'@context': 'http://schema.org', '@type': 'Organization', 'logo': 'https://www.vitenda.de/documents/logo/logo_vitenda_02_646.png', 'url': 'https://www.vitenda.de'}, {'@context': 'http://schema.org/', '@type': 'BreadcrumbList', 'itemListElement': [{'@type': 'ListItem', 'item': {'@id': 'https://www.vitenda.de/search', 'name': 'Artikelsuche'}, 'position': 1}, {'@type': 'ListItem', 'item': {'@id': '', 'name': 'Ihre Suchergebnisse für ' "<b>'11287708'</b> (1 " 'Produkte)'}, 'position': 2}]}, {'@context': 'http://schema.org/', '@type': 'Product', 'brand': {'@type': 'Organization', 'name': 'ALIUD Pharma GmbH'}, 'description': '', 'gtin': '', 'image': 'https://cdn1.apopixx.de/300/web_schraeg_png/11287708.png?ver=1649058520', 'itemCondition': 'https://schema.org/NewCondition', 'name': 'GINKGO AL 240 mg Filmtabletten', 'offers': {'@type': 'Offer', 'availability': 'http://schema.org/InStock', 'deliveryLeadTime': {'@type': 'QuantitativeValue', 'minValue': '3'}, 'price': 96.36, 'priceCurrency': 'EUR', 'priceValidUntil': '19-06-2022 18:41:54', 'url': 'https://www.vitenda.de/ginkgo-al-240-mg-filmtabletten.11287708'}, 'productID': '11287708', 'sku': '11287708', 'url': 'https://www.vitenda.de/ginkgo-al-240-mg-filmtabletten.11287708'}]} # 初始化结果存储 product_info = None # 遍历所有json-ld条目 for item in data.get('json-ld', []): # 跳过没有@type字段的无效条目 if '@type' not in item: continue # 匹配Product类型,兼容@type为数组的特殊情况 item_type = item['@type'] if (isinstance(item_type, str) and item_type == 'Product') or \ (isinstance(item_type, list) and 'Product' in item_type): # 提取目标字段 offers = item.get('offers', {}) product_info = { 'product_name': item.get('name'), 'url': item.get('url', offers.get('url')), 'price': offers.get('price'), 'price_currency': offers.get('priceCurrency'), 'availability': offers.get('availability') } # 如果只需要第一个匹配的Product可以加break终止遍历 break # 输出结果 if product_info: print("找到Product类型条目,提取信息如下:") for k, v in product_info.items(): print(f"{k}: {v}") else: print("元数据中未找到Product类型条目")
注意事项
- 原测试代码的判断逻辑
if '@context' in data['json-ld'][0]['@context']是无效的:该逻辑是判断字符串@context是否存在于第一个条目的@context属性值(也就是schema的url字符串)里,和查找Product条目的需求完全无关。 - 硬编码下标取
data['json-ld'][3]的写法兼容性很差,不同页面的JSON-LD条目顺序可能变化,遍历匹配的方式更稳定。 - 部分场景下
offers字段可能是数组(对应同一个产品的多个销售渠道/规格报价),如果遇到这类数据可以再加一层遍历提取所有报价信息即可。 - 库存字段
availability的返回值是schema定义的固定路径,比如InStock代表有货、OutOfStock代表缺货,可以根据字符串末尾的类型值做状态判断。
内容的提问来源于stack exchange,提问作者merlin
相关产品推荐
相关产品推荐

