You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用BeautifulSoup爬取时递增URL中的数字(如youtube.com/user/1/→2/)

基于BeautifulSoup实现循环递增URL爬取YouTube用户页面

以下是直接可用的实现方案,核心通过循环自增数字构造目标URL,结合BeautifulSoup解析页面内容:

所需依赖

先安装必要的库:

pip install requests beautifulsoup4

基础实现代码(指定爬取范围)

import requests
from bs4 import BeautifulSoup
import time

# 起始与终止用户ID
start_id = 1
end_id = 10  # 示例:爬取ID1到ID10的页面

for user_id in range(start_id, end_id + 1):
    # 构造目标URL
    url = f"https://youtube.com/user/{user_id}/"
    
    try:
        # 模拟浏览器请求,避免被反爬拦截
        headers = {
            "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
        }
        response = requests.get(url, headers=headers)
        response.raise_for_status()  # 抛出HTTP请求异常
        
        # 解析页面内容
        soup = BeautifulSoup(response.text, "html.parser")
        
        # 示例:提取页面标题,可替换为你需要的内容逻辑
        page_title = soup.title.string if soup.title else "无标题"
        print(f"用户ID {user_id} | 页面标题:{page_title}")
        
        # 这里可以添加其他爬取逻辑,比如提取用户简介、视频列表等
        # 示例:提取用户简介
        # profile_desc = soup.find("meta", property="og:description")["content"] if soup.find("meta", property="og:description") else "无简介"
        
        # 控制请求频率,避免触发反爬
        time.sleep(2)
        
    except requests.exceptions.HTTPError as e:
        print(f"用户ID {user_id} 请求失败:{e}")
    except Exception as e:
        print(f"用户ID {user_id} 处理出错:{str(e)}")

进阶实现(无限循环直到连续失败)

如果不需要指定终止ID,而是遇到连续多次无效页面后自动停止,可以用这个版本:

import requests
from bs4 import BeautifulSoup
import time

user_id = 1
failed_count = 0
max_failed = 3  # 连续3次失败就停止爬取

while failed_count < max_failed:
    url = f"https://youtube.com/user/{user_id}/"
    
    try:
        headers = {
            "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
        }
        response = requests.get(url, headers=headers)
        response.raise_for_status()
        
        soup = BeautifulSoup(response.text, "html.parser")
        page_title = soup.title.string if soup.title else "无标题"
        print(f"用户ID {user_id} | 页面标题:{page_title}")
        
        failed_count = 0  # 成功爬取后重置失败计数器
        time.sleep(2)
        
    except requests.exceptions.HTTPError:
        print(f"用户ID {user_id} 页面不存在或请求被拒绝")
        failed_count += 1
    except Exception as e:
        print(f"用户ID {user_id} 处理出错:{str(e)}")
        failed_count += 1
        
    user_id += 1

print(f"连续{max_failed}次请求失败,停止爬取")

关键注意点

  • 反爬规避:必须添加User-Agent模拟浏览器,同时通过time.sleep()控制请求间隔,否则很容易被YouTube封禁IP。
  • 异常处理:捕获HTTP错误(如404)和未知异常,保证循环不会因为单个页面出错而中断。
  • 内容提取:根据需求修改soup解析逻辑,BeautifulSoup的find()、find_all()方法可以满足大部分页面元素提取需求。

内容的提问来源于stack exchange,提问作者user19476773

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 18:01:14