使用BeautifulSoup爬取Realtor页面仅返回首个数据点的问题
问题:爬取Realtor网站休斯顿地区房产经纪人数据仅获取首个条目,如何获取每页全部20条数据?
用户提供的代码用于爬取Realtor网站休斯顿地区的房产经纪人姓名与电话数据,每页应返回20条数据,但当前仅能获取首个数据点。页面结构显示每页20个经纪人条目、姓名及电话均使用相同类名。
原代码如下:
from bs4 import BeautifulSoup import requests import xlsxwriter number = 1 counter = 0 row = 0 wb = xlsxwriter.Workbook('try002.xlsx') sheet = wb.add_worksheet() while number <= 150: url = "https://www.realtor.com/realestateagents/houston_tx/photo-1/sort-activelistings/pg-{}".format(number) base = 'https://www.realtor.com' response = requests.get(url) soup = BeautifulSoup(response.content, 'html.parser') container = soup.find_all('ul', class_='jsx-1526930885') # print(container) for con in container: try: name = con.find('div', class_='jsx-2987058905 agent-name text-bold').text except: name = 'no price found' try: pnumber = con.find('div', class_='jsx-2987058905 agent-phone hidden-xs hidden-xxs').text except: pnumber = 'no type found' sheet.write(row, 0, name) sheet.write(row, 1, pnumber) counter += 1 row += 1 print(counter) number += 1 print(url) wb.close()
问题原因
原代码中soup.find_all('ul', class_='jsx-1526930885')获取的是整个经纪人列表的容器(ul元素),而非单个经纪人条目。遍历这个容器时,con.find(...)只会提取容器内第一个匹配的姓名和电话,因此仅能得到1条数据。
修改方案
需要定位到每个经纪人对应的li条目,而非整个ul容器。根据页面结构,每个经纪人条目是ul下的li元素,修改代码如下:
from bs4 import BeautifulSoup import requests import xlsxwriter number = 1 counter = 0 row = 0 wb = xlsxwriter.Workbook('try002.xlsx') sheet = wb.add_worksheet() while number <= 150: url = "https://www.realtor.com/realestateagents/houston_tx/photo-1/sort-activelistings/pg-{}".format(number) response = requests.get(url) soup = BeautifulSoup(response.content, 'html.parser') # 定位到每个经纪人的li条目 agent_items = soup.find_all('li', class_='jsx-1526930885 agent-list-item') for item in agent_items: try: name = item.find('div', class_='jsx-2987058905 agent-name text-bold').text.strip() except: name = '姓名未找到' try: pnumber = item.find('div', class_='jsx-2987058905 agent-phone hidden-xs hidden-xxs').text.strip() except: pnumber = '电话未找到' sheet.write(row, 0, name) sheet.write(row, 1, pnumber) counter += 1 row += 1 print(counter) number += 1 print(f"已爬取页面: {url}") wb.close()
修改说明
- 定位单个经纪人条目:将
find_all('ul', ...)改为find_all('li', class_='jsx-1526930885 agent-list-item'),直接获取每页所有20个经纪人的li元素。 - 遍历每个条目提取数据:循环遍历每个li元素,从每个条目中单独提取姓名和电话,确保每条数据对应一个经纪人。
- 优化异常提示:将错误提示改为更贴合场景的“姓名未找到”“电话未找到”,并使用
.strip()去除文本中的多余空格。
内容的提问来源于stack exchange,提问作者eInvincible
相关产品推荐
相关产品推荐

