You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页爬取无法获取多页:BoxOffice网站多页爬取仅获单页问题排查

BoxOffice Mojo多页票房爬取循环问题排查与修复

原代码存在的关键问题

  • URL语法错误:字符串拼接时缺少右括号,导致请求无法正常发起
  • 终止条件错误:对整数page_number使用len()函数,逻辑完全错误,应直接判断数值是否达到阈值
  • 页码更新错误:page_number =+ 200是将变量赋值为200,而非累加200,导致页码无法推进,循环卡在同一页
  • 数据存储错误:All_movies = []放在循环内部,每次循环都会清空列表,最终只能保留最后一页数据
  • 异常捕获过于宽泛:无差别break会掩盖网络请求、解析等异常,无法定位问题
  • 数据行选择错误:soup.find_all("tr")会包含表头行,导致无效数据被加入列表

修正后的代码

import requests
from bs4 import BeautifulSoup as bs

# 初始化存储所有电影数据的列表,放在循环外
All_movies = []
page_number = 0  # 初始偏移量从0开始,对应第一页
max_offset = 600  # 设置最大偏移量,可根据实际需求调整

while True:
    try:
        # 修正URL拼接的括号问题
        url = f'https://www.boxofficemojo.com/chart/ww_top_lifetime_gross/?area=XWW&offset={page_number}'
        response = requests.get(url)
        # 检查请求是否成功,避免无效响应
        response.raise_for_status()
        soup = bs(response.text, 'html.parser')
        
        # 修正终止条件:判断页码偏移量是否达到阈值
        if page_number >= max_offset:
            print("爬取完成")
            break
        
        # 只选择tbody内的tr标签,过滤表头行
        detail = soup.select("tbody tr")
        for de in detail:
            title_elem = de.find("td", {"class": "a-text-left mojo-field-type-title"})
            gross_elem = de.find("td", {"class": "a-text-right mojo-field-type-money"})
            year_elem = de.find("td", {"class": "a-text-left mojo-field-type-year"})
            
            title = title_elem.text.strip() if title_elem else None
            # 清理票房格式,去掉$和逗号便于后续数值处理
            worldwide_gross = gross_elem.text.strip().replace('$', '').replace(',', '') if gross_elem else None
            year = year_elem.text.strip() if year_elem else None
            
            # 只保留有完整数据的条目
            if title and worldwide_gross and year:
                All_movies.append([title, worldwide_gross, year])
        
        # 正确累加页码偏移量
        page_number += 200
        print(f"已爬取偏移量为{page_number - 200}的页面,URL: {url}")
        
    except requests.exceptions.RequestException as e:
        print(f"请求出错: {e}")
        break
    except Exception as e:
        print(f"解析出错: {e}")
        break

# 可选:输出爬取结果统计
print(f"共爬取{len(All_movies)}条有效电影数据")

额外说明

  • 初始偏移量改为0,因为BoxOffice Mojo第一页的offset参数值为0,原代码从200开始会跳过第一页数据
  • 新增response.raise_for_status(),用于捕获HTTP请求错误(如404、500状态码)
  • 优化了票房数据的格式处理,去除多余符号便于后续分析
  • 拆分异常捕获逻辑,分别处理请求和解析阶段的错误,方便排查问题

内容的提问来源于stack exchange,提问作者Khaled Hamdy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 04:35:02