You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

无页码URL下如何用Python Requests+BeautifulSoup爬取100页链接?

解决方案

针对你的问题,无页码式的翻页爬取通常有两种实现方式,结合你的场景具体说明:

1. 查找页面内的「下一页」链接

很多无页码的网站会在页面底部提供「下一页」按钮,你可以通过解析页面找到该按钮的跳转链接,循环请求直到获取100页内容:

修改后的代码示例

import requests
from bs4 import BeautifulSoup
import time

links = []
current_url = 'https://unicorner.news'
page_count = 0

# 循环爬取最多100页
while current_url and page_count < 100:
    try:
        # 发送请求,加入延时避免反爬
        r = requests.get(current_url)
        r.raise_for_status()
        soup = BeautifulSoup(r.text, "html.parser")
        
        # 提取当前页的文章链接
        current_page_links = soup.find_all('a', class_='utils_postLink__m_2J6')
        for link in current_page_links:
            links.append(link['href'])
        
        # 查找下一页链接(需根据页面实际元素调整class或文本)
        next_page_btn = soup.find('a', text='下一页')  # 或替换为实际的class,比如'pagination-next'
        if next_page_btn:
            current_url = next_page_btn['href']
            # 处理相对路径,拼接为绝对URL
            if not current_url.startswith('http'):
                current_url = 'https://unicorner.news' + current_url
        else:
            # 没有下一页则终止循环
            break
        
        page_count += 1
        time.sleep(1)  # 每秒请求一次,降低被封风险
        
    except requests.exceptions.RequestException as e:
        print(f"请求出错: {e}")
        break

print(f"共获取到 {len(links)} 条链接")

2. 模拟AJAX滚动加载请求

如果网站是通过滚动页面触发AJAX请求加载新内容,你需要先通过浏览器开发者工具(F12 → Network标签)找到加载更多内容的API接口:

  • 滚动页面,观察XHR类型的请求,找到返回文章列表的接口
  • 分析接口参数(通常是offset偏移量或page页码),比如offset=10表示从第11条开始加载

代码示例(假设API接口为示例格式)

import requests
import time

links = []
page_size = 10  # 默认每页10条
total_pages = 100

for page in range(total_pages):
    offset = page * page_size
    # 替换为你找到的实际API接口
    api_url = f'https://unicorner.news/api/posts?offset={offset}&limit={page_size}'
    
    try:
        r = requests.get(api_url)
        r.raise_for_status()
        data = r.json()
        
        # 从返回的JSON中提取文章链接(需根据实际字段调整)
        for post in data.get('data', []):
            links.append(post['url'])
        
        time.sleep(1)
        
    except requests.exceptions.RequestException as e:
        print(f"请求API出错: {e}")
        break

print(f"共获取到 {len(links)} 条链接")

关于「设置每页显示100条」的问题

这取决于网站是否提供该功能:

  • 如果页面上有「每页显示条数」的下拉选项,你可以先提交修改条数的请求(通常是表单提交或修改URL参数),再按上述方式爬取
  • 如果网站没有提供该选项,则无法强制设置,只能按默认每页10条爬取100页

内容的提问来源于stack exchange,提问作者jpy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 10:45:32