You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在while True循环中添加条件限制Python爬虫的爬取页数

爬虫指定页数爬取修改方案

修改逻辑

  • 新增自定义目标页数变量,可自由设置你需要爬取的页数(10/20/100均可)
  • 新增已爬页数计数器,初始值设为1(对应你代码中提前请求的第一页内容)
  • 将原无限循环while True替换为计数器判断条件,每完成一页爬取计数器自增1
  • 额外优化:新增全局空DataFrame拼接所有页面数据,解决原代码仅保留最后一页数据的问题;新增下一页存在性判断,避免目标页数超过网站实际总页数时报错

修改后完整代码

import requests 
from bs4 import BeautifulSoup
import pandas as pd

# ========== 自定义参数区 ==========
# 只需修改这个值即可调整爬取总页数,示例为爬取10页
TARGET_PAGES = 10

# 初始化变量
current_page = 1
all_data = pd.DataFrame()

re=requests.get("https://katmoviehd.sk/")
soup=BeautifulSoup(re.text,"html.parser")

while current_page <= TARGET_PAGES:
    page = soup.find_all('h2')[1:]
    
    Category = soup.find_all('span', class_ = 'meta-category')
    Category_list = []
    for i in Category:
        Category2 = i.text
        Category_list.append(Category2)
    
    link_list = []
    for i in page:
        link = i.find("a")['href']
        link_list.append(link)
        
    title_list = []    
    for i in page:
        title = i.find("a")['title']
        title_list.append(title)
        
    # 拼接当前页数据到全局数据集
    Table = pd.DataFrame({'Links':link_list, 'Title':title_list, 'Category':Category_list})
    all_data = pd.concat([all_data, Table], ignore_index=True)
   
    # 判断是否还有下一页、是否达到目标爬取页数
    next_page_tag = soup.find('a', class_ = 'next page-numbers')
    if not next_page_tag or current_page >= TARGET_PAGES:
        break
    next_page = next_page_tag.get('href')
    
    # 请求下一页内容
    url = next_page
    page = requests.get(url)
    soup = BeautifulSoup(page.text, 'lxml')
    current_page += 1

# 爬取完成后可以导出所有数据,示例导出为csv文件
# all_data.to_csv("爬取结果.csv", index=False, encoding='utf-8-sig')

使用说明

你只需修改代码开头TARGET_PAGES的赋值即可调整爬取页数,比如需要爬取100页就改为TARGET_PAGES = 100。爬取完成后所有页面的数据都存储在all_data变量中,你可以根据需要做后续处理。

内容的提问来源于stack exchange,提问作者Shovo Murad

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 15:27:02