You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python提取LD+JSON中的图片值

提取LD+JSON中的图片值的正确方法

你的代码存在几个细节问题,修正后即可正确提取图片值,以下是调整后的方案:

代码问题修正

  1. 模块导入错误:import request 应改为 import requests
  2. 缩进错误:headers 定义的缩进需调整,与其他代码保持同一层级
  3. 请求未携带Headers:直接请求可能被网站反爬机制拦截,需在requests.get()中传入headers参数
  4. LD+JSON结构可能嵌套:并非所有返回的JSON对象都直接包含images字段,需针对酒店类Schema的结构进行解析

修正后的代码

import json
import requests
from bs4 import BeautifulSoup

url = 'https://www.tiket.com/hotel/indonesia/ideas-hotel-bandung-108001534490330380?checkin=2023-01-26&checkout=2023-01-27&room=1&adult=1&soldOut='

# 设置请求头模拟浏览器访问
headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/96.0.4664.110 Safari/537.36'}

# 发送请求并解析页面
response = requests.get(url, headers=headers)
soup = BeautifulSoup(response.content, 'html.parser')

# 提取所有非空的LD+JSON数据
ld_json_list = [json.loads(script.string) for script in soup.find_all('script', type='application/ld+json') if script.string]

# 遍历解析每个LD+JSON对象,提取图片
for item in ld_json_list:
    if isinstance(item, dict):
        # 检查是否有直接的images字段
        images = item.get('images')
        if images:
            print("提取到图片链接:")
            for img in images:
                print(img)
        # 检查嵌套的mainEntity字段(酒店页面常见结构)
        main_entity = item.get('mainEntity')
        if main_entity and isinstance(main_entity, dict):
            entity_images = main_entity.get('images')
            if entity_images:
                print("\n从mainEntity提取到图片链接:")
                for img in entity_images:
                    print(img)

关键说明

  • 加入if script.string过滤空script标签,避免解析报错
  • 针对酒店页面Schema的常见嵌套结构,增加了对mainEntity字段的检查,确保不会遗漏嵌套的图片数据
  • 分来源打印图片链接,方便你直观查看数据结构

内容的提问来源于stack exchange,提问作者Si Doel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 13:05:22