You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解决Python+BeautifulSoup爬取OLX尼日利亚站仅获前20商品的问题?

解决OLX尼日利亚站爬虫分页问题:获取全部商品数据

嘿,我来帮你搞定这个分页爬取的问题!你现在遇到的情况太常见了——OLX这类电商网站默认只加载第一页的20个商品,要拿到全站(或目标分类下)的所有数据,得跟着它的分页规则来循环请求每一页。下面一步步给你讲清楚怎么改代码:

1. 先搞懂OLX的分页规则

打开OLX尼日利亚站,翻到第二页你会发现URL变成了 https://www.olx.com.ng/?page=2,第三页就是 ?page=3,以此类推。也就是说,我们只需要给基础URL加上page参数,就能构造出每一页的请求地址。

2. 确定何时停止爬取

总不能无限循环下去对吧?推荐用简单直接的判断方式:检查当前页面的商品列表是否为空,如果没有商品了,说明已经爬完所有页。

3. 必须加的反爬措施

OLX有反爬机制,直接用urlopen裸请求很容易被封。所以我们要:

  • 模拟浏览器的请求头(User-Agent)
  • 每次请求后加随机延迟,避免请求太频繁

修改后的完整代码

下面是调整后的代码,我加了详细注释,你可以直接用:

from bs4 import BeautifulSoup
from urllib.request import urlopen, Request
import csv
import random
import time

# 模拟浏览器请求头,避免被识别为爬虫
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
}

def scrape_all_olx_products():
    # 打开CSV文件,写入表头
    with open('olx_full_products.csv', 'w', newline='', encoding='utf-8') as csv_file:
        writer = csv.writer(csv_file)
        writer.writerow(['商品标题', '商品价格', '简短描述', '商品链接'])
        
        page_number = 1
        while True:
            # 构造当前页的URL
            current_url = f'https://www.olx.com.ng/?page={page_number}'
            print(f"正在爬取第 {page_number} 页...")
            
            try:
                # 带请求头发起请求
                request = Request(current_url, headers=headers)
                response = urlopen(request)
                soup = BeautifulSoup(response, 'html.parser')
                
                # 定位商品列表(注意:OLX的class可能会更新,建议用F12检查当前页面的商品容器class)
                product_items = soup.find_all('div', class_='ef447dde')
                
                # 如果没有商品,说明到最后一页了,停止循环
                if not product_items:
                    print("所有页面爬取完成!")
                    break
                
                # 遍历每个商品,提取信息
                for item in product_items:
                    # 提取标题
                    title = item.find('h6', class_='a5112ca8').get_text(strip=True) if item.find('h6', class_='a5112ca8') else '无标题'
                    # 提取价格
                    price = item.find('span', class_='_9a4e3964').get_text(strip=True) if item.find('span', class_='_9a4e3964') else '无价格'
                    # 提取简短描述
                    description = item.find('span', class_='_205654f7').get_text(strip=True) if item.find('span', class_='_205654f7') else '无描述'
                    # 提取商品链接
                    product_link = 'https://www.olx.com.ng' + item.find('a')['href'] if item.find('a') else '无链接'
                    
                    # 写入CSV
                    writer.writerow([title, price, description, product_link])
                
                # 随机延迟1-3秒,降低被封风险
                sleep_time = random.uniform(1, 3)
                time.sleep(sleep_time)
                
                # 页码加1,继续下一页
                page_number += 1
                
            except Exception as e:
                print(f"爬取第 {page_number} 页时出错: {str(e)}")
                # 出错时停止循环,你也可以改成重试逻辑
                break

if __name__ == '__main__':
    scrape_all_olx_products()

几个重要的注意事项

  • 元素class可能变化:OLX偶尔会更新页面结构,如果你发现数据提取失败,记得用浏览器开发者工具(F12)重新查看商品元素的class,替换代码里对应的class值。
  • 分类爬取:如果你想爬特定分类(比如手机、家电),只需要把基础URL换成分类页的地址,比如手机分类的URL是https://www.olx.com.ng/mobile-phones/,分页规则还是一样的,构造https://www.olx.com.ng/mobile-phones/?page={page_number}即可。
  • IP封禁风险:如果爬取的页数很多,建议使用代理IP,或者增加延迟时间,避免被OLX封禁IP。

内容的提问来源于stack exchange,提问作者ade desmond

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 06:58:08