You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python网络爬虫:如何从房产网页HTML提取公寓分段信息存入表格

解决方案

核心逻辑:你之前直接提取整个<a>标签文本会丢失结构,正确做法是通过每个字段绑定的固定CSS类名定位提取,不受特征数量、排序变化的影响。

实现代码

import requests
from bs4 import BeautifulSoup
import re
import pandas as pd

# 原有页面获取逻辑
base = 'https://www.booli.se/'
addition = 'kungsholmen,vasastan,gardet/115353,115349,115347'
url = base+addition
notar = requests.get(url)
page = notar.content
soup = BeautifulSoup(page, 'lxml')
links = soup.find_all('a', href=re.compile('/annons/'))

# 初始化列表存储所有结构化房源数据
apt_list = []

for row in links:
    # 初始化单条房源信息字典
    apt_info = {}
    
    # 1. 房源详情页完整链接
    apt_info['详情链接'] = base + row['href'].lstrip('/')
    
    # 2. 房源特征(兼容0-5个的浮动数量)
    features = row.find_all('p', class_='_1Csan')
    apt_info['房源特征'] = [f.get_text(strip=True) for f in features] if features else []
    
    # 3. 公寓名称
    name_tag = row.find('h3', class_='_3R27q')
    apt_info['公寓名称'] = name_tag.get_text(strip=True) if name_tag else None
    
    # 4. 售价
    price_tag = row.find('p', class_='_3HprD')
    apt_info['售价'] = price_tag.get_text(strip=True) if price_tag else None
    
    # 5. 估值
    estimate_tag = row.find('div', class_='MsC3E')
    apt_info['估值'] = estimate_tag.get_text(strip=True) if estimate_tag else None
    
    # 6. 基础信息块(区域、户型、面积、月费等)
    base_info_block = row.find('div', class_='_3IN5U')
    if base_info_block:
        base_items = [p.get_text(strip=True) for p in base_info_block.find_all('p')]
        # 按顺序赋值,自动兼容字段缺失
        apt_info['区域'] = base_items[0] if len(base_items)>=1 else None
        apt_info['房屋类型'] = base_items[1] if len(base_items)>=2 else None
        apt_info['房间数'] = base_items[2] if len(base_items)>=3 else None
        apt_info['面积'] = base_items[3] if len(base_items)>=4 else None
        apt_info['月费'] = base_items[4] if len(base_items)>=5 else None
    
    apt_list.append(apt_info)

# 转换为结构化表格,可直接导出为CSV/Excel
df = pd.DataFrame(apt_list)
# 导出CSV示例
df.to_csv('booli房源数据.csv', index=False, encoding='utf-8-sig')

方案说明

  • 所有字段通过固定CSS类名定位,只要网站前端没有改版修改类名,就可以稳定提取,不受单房源特征缺失、排序变化的影响
  • 提取时做了缺失值兼容,没有对应字段的房源会自动填充None,不会报错中断爬取
  • 最终输出的DataFrame可以直接导出为CSV、Excel等结构化格式,满足存储需求

内容的提问来源于stack exchange,提问作者karwi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.02 08:57:04