You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup爬取Realtor页面仅返回首个数据点的问题

问题:爬取Realtor网站休斯顿地区房产经纪人数据仅获取首个条目,如何获取每页全部20条数据?

用户提供的代码用于爬取Realtor网站休斯顿地区的房产经纪人姓名与电话数据,每页应返回20条数据,但当前仅能获取首个数据点。页面结构显示每页20个经纪人条目、姓名及电话均使用相同类名。

原代码如下:

from bs4 import BeautifulSoup
import requests
import xlsxwriter

number = 1
counter = 0
row = 0
wb = xlsxwriter.Workbook('try002.xlsx')
sheet = wb.add_worksheet()
while number <= 150:
    url = "https://www.realtor.com/realestateagents/houston_tx/photo-1/sort-activelistings/pg-{}".format(number)
    base = 'https://www.realtor.com'
    response = requests.get(url)
    soup = BeautifulSoup(response.content, 'html.parser')
    container = soup.find_all('ul', class_='jsx-1526930885')
    # print(container)
    for con in container:
        try:
            name = con.find('div', class_='jsx-2987058905 agent-name text-bold').text
        except:
            name = 'no price found'
        try:
            pnumber = con.find('div', class_='jsx-2987058905 agent-phone hidden-xs hidden-xxs').text
        except:
            pnumber = 'no type found'
        sheet.write(row, 0, name)
        sheet.write(row, 1, pnumber)
        counter += 1
        row += 1
        print(counter)
    number += 1
    print(url)
wb.close()

问题原因

原代码中soup.find_all('ul', class_='jsx-1526930885')获取的是整个经纪人列表的容器(ul元素),而非单个经纪人条目。遍历这个容器时,con.find(...)只会提取容器内第一个匹配的姓名和电话,因此仅能得到1条数据。

修改方案

需要定位到每个经纪人对应的li条目,而非整个ul容器。根据页面结构,每个经纪人条目是ul下的li元素,修改代码如下:

from bs4 import BeautifulSoup
import requests
import xlsxwriter

number = 1
counter = 0
row = 0
wb = xlsxwriter.Workbook('try002.xlsx')
sheet = wb.add_worksheet()
while number <= 150:
    url = "https://www.realtor.com/realestateagents/houston_tx/photo-1/sort-activelistings/pg-{}".format(number)
    response = requests.get(url)
    soup = BeautifulSoup(response.content, 'html.parser')
    # 定位到每个经纪人的li条目
    agent_items = soup.find_all('li', class_='jsx-1526930885 agent-list-item')
    for item in agent_items:
        try:
            name = item.find('div', class_='jsx-2987058905 agent-name text-bold').text.strip()
        except:
            name = '姓名未找到'
        try:
            pnumber = item.find('div', class_='jsx-2987058905 agent-phone hidden-xs hidden-xxs').text.strip()
        except:
            pnumber = '电话未找到'
        sheet.write(row, 0, name)
        sheet.write(row, 1, pnumber)
        counter += 1
        row += 1
        print(counter)
    number += 1
    print(f"已爬取页面: {url}")
wb.close()

修改说明

  1. 定位单个经纪人条目:将find_all('ul', ...)改为find_all('li', class_='jsx-1526930885 agent-list-item'),直接获取每页所有20个经纪人的li元素。
  2. 遍历每个条目提取数据:循环遍历每个li元素,从每个条目中单独提取姓名和电话,确保每条数据对应一个经纪人。
  3. 优化异常提示:将错误提示改为更贴合场景的“姓名未找到”“电话未找到”,并使用.strip()去除文本中的多余空格。

内容的提问来源于stack exchange,提问作者eInvincible

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 22:07:28