求助:使用BeautifulSoup爬取公寓网站时嵌套循环返回重复信息
问题分析与解决
核心问题
你的代码使用三层嵌套循环,导致笛卡尔积式的重复输出:
- 外层循环取第一个
placard-content的价格和户型后,会遍历所有地址、所有公寓名称,最终所有结果都复用了第一个户型、价格和第一个地址,仅名称变化。 - 根本错误是未将每个公寓的信息绑定到其独立卡片容器,而是全局提取所有地址、名称后强行组合,无法实现信息一一对应。
修正方案
找到每个公寓的独立外层容器(li.mortar-wrapper),在容器内部提取对应信息,确保信息匹配:
import requests from bs4 import BeautifulSoup url = "https://www.apartments.com/atlanta-ga/?bb=lnwszyjy-H4lu8uqH" header = {'User-Agent': 'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/79.0.3945.88 Safari/537.36'} page = requests.get(url, headers=header) soup = BeautifulSoup(page.content, 'html.parser') # 遍历每个独立的公寓卡片容器 for apartment in soup.find_all('li', class_="mortar-wrapper"): # 从当前卡片提取各字段 name = apartment.find('span', class_="js-placardTitle title") location = apartment.find('div', class_="property-address js-url") beds = apartment.find('p', class_="property-beds") price = apartment.find('p', class_="property-pricing") # 处理可能的空值,避免元素缺失报错 if all([name, location, beds, price]): info = [ name.get_text(strip=True), location.get_text(strip=True), beds.get_text(strip=True), price.get_text(strip=True) ] print(info)
补充说明
- 昨天能正常爬取是巧合:当时网站DOM结构中,全局提取的
lists、addresses、properties数量刚好一致,循环时能勉强对应,但这种写法依赖网站结构稳定性,极易失效。 get_text(strip=True)可自动去除文本前后空格、换行,优化输出格式。- 空值判断避免了部分公寓卡片信息缺失时,代码直接中断的问题。
内容的提问来源于stack exchange,提问作者user15564501
相关产品推荐
相关产品推荐

