使用BeautifulSoup爬取汽车网站弹窗内卖家手机号的问题
解决弹窗内卖家手机号的爬取问题
首先,你遇到的情况很典型:弹窗里的手机号不会直接出现在页面初始HTML中,这类网站通常会在用户点击"显示手机号"按钮时,通过AJAX异步请求从后端接口拉取数据。针对auto.ria.com,我们可以按下面的思路处理:
核心步骤
- 从车辆详情页URL里提取车辆ID(比如
https://auto.ria.com/uk/auto_1234567.html里的1234567就是ID) - 找到网站获取手机号的API接口,用车辆ID构造请求地址
- 发送请求到API,解析返回的JSON数据提取手机号
修改后的代码
from bs4 import BeautifulSoup import requests import json import time page = 0 # 模拟浏览器请求头,避免被反爬拦截 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } while page <= 500: url = f"https://auto.ria.com/uk/search/?indexName=auto,order_auto,newauto_search&categories.main.id=1&country.import.usa.not=-1&price.USD.gte=1111&price.currency=1&abroad.not=0&custom.not=1&page={page}" response = requests.get(url, headers=headers) soup = BeautifulSoup(response.text, "html.parser") cars = soup.find_all("a", class_="m-link-ticket") if cars: for link in cars: car_url = link.get('href') print(f"车辆链接: {car_url}") # 从链接中提取车辆ID car_id = car_url.split("_")[-1].replace(".html", "") # 构造手机号API请求地址 phone_api_url = f"https://auto.ria.com/users/phones/{car_id}" # 添加延迟,避免请求过于频繁被封 time.sleep(1) phone_response = requests.get(phone_api_url, headers=headers) if phone_response.status_code == 200: try: phone_data = json.loads(phone_response.text) # 提取手机号(根据API返回结构调整) if "phones" in phone_data: for item in phone_data["phones"]: print(f"卖家手机号: {item['phone']}") except json.JSONDecodeError: print("手机号数据解析失败") else: print("获取手机号请求失败") page += 1 else: break
关键注意点
- 反爬规避:一定要加请求头和请求延迟,频繁请求很容易被网站封禁IP
- API变动:如果后续网站修改了接口,你可以用浏览器开发者工具(F12)抓包,查看点击"显示手机号"时发送的真实请求地址和参数
- 数据结构:如果API返回的JSON结构变化,需要对应调整代码里的字段提取逻辑
内容的提问来源于stack exchange,提问作者EzNo
相关产品推荐
相关产品推荐

