You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup从Letterboxd抓取1980年代电影列表?

修复Letterboxd电影标题抓取为空的问题

核心问题分析

  1. 反爬拦截:直接用requests.get发起请求会被网站识别为非浏览器请求,返回的页面不含电影列表内容。
  2. 元素定位错误:原代码依赖的元素class已失效,网站页面结构有更新。

修复后的代码

import requests
from bs4 import BeautifulSoup

url = "https://letterboxd.com/films/decade/1980s/"
# 模拟浏览器请求头,避免被反爬拦截
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
}

response = requests.get(url, headers=headers)

if response.status_code == 200:
    soup = BeautifulSoup(response.text, "html.parser")
    # 定位正确的电影列表元素
    film_elements = soup.find_all("li", class_="poster-container")
    film_titles = []
    for item in film_elements:
        # 从a标签的title属性提取电影名称(更稳定)
        film_title = item.find("a")["title"]
        film_titles.append(film_title)
    print(film_titles)
else:
    print(f"页面请求失败,状态码:{response.status_code}")

关键修改点

  • 添加请求头:通过User-Agent模拟浏览器访问,绕过基础反爬检测。
  • 修正元素定位:
    • 电影列表的li元素class改为poster-container(原代码的listitem poster-container已不存在)。
    • 从a标签的title属性提取标题,比依赖span标签更稳定,避免页面结构小变动导致失效。

内容的提问来源于stack exchange,提问作者Dominik Volk

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 12:46:14