如何从Letterboxd的application/ld+json脚本提取海报图片链接?
解决Letterboxd海报抓取的脚本内容解析问题
首先,你原代码里的string="image"是精确匹配,几乎找不到目标标签,得改成模糊匹配包含image关键词的script:
response = page.text soup = bs(response, 'html.parser') # 匹配包含"image"的application/ld+json类型script scripts = soup.find_all("script", {"type": "application/ld+json"}, string=lambda t: t and 'image' in t)
接下来处理CDATA:Letterboxd的这类script内容通常被CDATA包裹,但BeautifulSoup会自动暴露CDATA里的文本,直接用.string就能拿到原始JSON字符串,不需要额外处理CDATA标签。
然后解析JSON并提取图片链接:
import json for script in scripts: try: # 提取script内的文本,自动跳过CDATA标记 json_data = json.loads(script.string) # 兼容单个对象或数组形式的结构化数据 if isinstance(json_data, list): json_data = json_data[0] # 处理image字段的两种格式:直接字符串或带url的对象 if isinstance(json_data.get('image'), dict): poster_url = json_data['image'].get('url') else: poster_url = json_data.get('image') if poster_url: print("海报链接:", poster_url) except json.JSONDecodeError: print("解析JSON失败,跳过该script")
关键说明
- 用
lambda匹配字符串是因为application/ld+json里是完整JSON结构,image只是其中一个字段,精确匹配根本找不到目标标签。 - 部分页面的结构化数据是数组格式,需要先判断类型取第一个元素。
- 部分场景下
image是包含url属性的对象,不是直接字符串,要做兼容处理。
内容的提问来源于stack exchange,提问作者Ansh Anand
相关产品推荐
相关产品推荐

