使用BeautifulSoup爬取房产网站时findAll无法采集多页数据问题求助
无法采集全量页面的原因
你的代码仅写死了第一页的请求URL,Zoopla的分页结果通过URL参数page_number控制分页,比如第2页的URL会追加&page_number=2,第N页对应&page_number=N,你没有循环遍历所有分页的URL发起请求,自然只能拿到第一页的数据。
另外补充两个注意点:
- 批量请求时建议给
requests.get加headers参数,携带正常浏览器的User-Agent,避免被网站反爬拦截 - 每次请求后要判断
response.status_code == 200再执行解析逻辑,避免异常报错
直接用BeautifulSoup方法提取价格的方案
完全可以,不需要把元素转成字符串做拆分操作,你拿到价格容器div后,直接提取内部p标签的文本内容即可,优化后的完整示例:
import requests from bs4 import BeautifulSoup as soup # 模拟浏览器请求头,避免反爬拦截 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } prices = [] # 遍历150页,也可以先抓取总页数动态设置循环上限更灵活 for page in range(1, 151): url = f'https://www.zoopla.co.uk/to-rent/property/central-london/?beds_max=5&price_frequency=per_month&q=Central%20London&results_sort=newest_listings&search_source=home&page_number={page}' response = requests.get(url, headers=headers) if response.status_code != 200: continue data = soup(response.content, 'lxml') for price_div in data.find_all('div', {'class': 'css-1e28vvi-PriceContainer e2uk8e7'}): # 直接提取p标签文本,再清洗数值即可 price_text = price_div.find('p').text.strip() price = int(price_text.split(' ')[0].replace('£', '').replace(',', '')) prices.append(price)
你之前把元素转成字符串拆分的方法稳定性很差,只要页面结构稍有变动就会报错,直接调用bs的节点选择方法容错性更高。
内容的提问来源于stack exchange,提问作者Angry Physicist
相关产品推荐
相关产品推荐

