You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从Letterboxd的application/ld+json脚本提取海报图片链接?

解决Letterboxd海报抓取的脚本内容解析问题

首先,你原代码里的string="image"是精确匹配,几乎找不到目标标签,得改成模糊匹配包含image关键词的script:

response = page.text
soup = bs(response, 'html.parser')
# 匹配包含"image"的application/ld+json类型script
scripts = soup.find_all("script", {"type": "application/ld+json"}, string=lambda t: t and 'image' in t)

接下来处理CDATA:Letterboxd的这类script内容通常被CDATA包裹,但BeautifulSoup会自动暴露CDATA里的文本,直接用.string就能拿到原始JSON字符串,不需要额外处理CDATA标签。

然后解析JSON并提取图片链接:

import json

for script in scripts:
    try:
        # 提取script内的文本,自动跳过CDATA标记
        json_data = json.loads(script.string)
        # 兼容单个对象或数组形式的结构化数据
        if isinstance(json_data, list):
            json_data = json_data[0]
        # 处理image字段的两种格式:直接字符串或带url的对象
        if isinstance(json_data.get('image'), dict):
            poster_url = json_data['image'].get('url')
        else:
            poster_url = json_data.get('image')
        if poster_url:
            print("海报链接:", poster_url)
    except json.JSONDecodeError:
        print("解析JSON失败,跳过该script")

关键说明

  • 用lambda匹配字符串是因为application/ld+json里是完整JSON结构,image只是其中一个字段,精确匹配根本找不到目标标签。
  • 部分页面的结构化数据是数组格式,需要先判断类型取第一个元素。
  • 部分场景下image是包含url属性的对象,不是直接字符串,要做兼容处理。

内容的提问来源于stack exchange,提问作者Ansh Anand

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 03:10:25