无页码URL下如何用Python Requests+BeautifulSoup爬取100页链接?
解决方案
针对你的问题,无页码式的翻页爬取通常有两种实现方式,结合你的场景具体说明:
1. 查找页面内的「下一页」链接
很多无页码的网站会在页面底部提供「下一页」按钮,你可以通过解析页面找到该按钮的跳转链接,循环请求直到获取100页内容:
修改后的代码示例
import requests from bs4 import BeautifulSoup import time links = [] current_url = 'https://unicorner.news' page_count = 0 # 循环爬取最多100页 while current_url and page_count < 100: try: # 发送请求,加入延时避免反爬 r = requests.get(current_url) r.raise_for_status() soup = BeautifulSoup(r.text, "html.parser") # 提取当前页的文章链接 current_page_links = soup.find_all('a', class_='utils_postLink__m_2J6') for link in current_page_links: links.append(link['href']) # 查找下一页链接(需根据页面实际元素调整class或文本) next_page_btn = soup.find('a', text='下一页') # 或替换为实际的class,比如'pagination-next' if next_page_btn: current_url = next_page_btn['href'] # 处理相对路径,拼接为绝对URL if not current_url.startswith('http'): current_url = 'https://unicorner.news' + current_url else: # 没有下一页则终止循环 break page_count += 1 time.sleep(1) # 每秒请求一次,降低被封风险 except requests.exceptions.RequestException as e: print(f"请求出错: {e}") break print(f"共获取到 {len(links)} 条链接")
2. 模拟AJAX滚动加载请求
如果网站是通过滚动页面触发AJAX请求加载新内容,你需要先通过浏览器开发者工具(F12 → Network标签)找到加载更多内容的API接口:
- 滚动页面,观察XHR类型的请求,找到返回文章列表的接口
- 分析接口参数(通常是
offset偏移量或page页码),比如offset=10表示从第11条开始加载
代码示例(假设API接口为示例格式)
import requests import time links = [] page_size = 10 # 默认每页10条 total_pages = 100 for page in range(total_pages): offset = page * page_size # 替换为你找到的实际API接口 api_url = f'https://unicorner.news/api/posts?offset={offset}&limit={page_size}' try: r = requests.get(api_url) r.raise_for_status() data = r.json() # 从返回的JSON中提取文章链接(需根据实际字段调整) for post in data.get('data', []): links.append(post['url']) time.sleep(1) except requests.exceptions.RequestException as e: print(f"请求API出错: {e}") break print(f"共获取到 {len(links)} 条链接")
关于「设置每页显示100条」的问题
这取决于网站是否提供该功能:
- 如果页面上有「每页显示条数」的下拉选项,你可以先提交修改条数的请求(通常是表单提交或修改URL参数),再按上述方式爬取
- 如果网站没有提供该选项,则无法强制设置,只能按默认每页10条爬取100页
内容的提问来源于stack exchange,提问作者jpy
相关产品推荐
相关产品推荐

