You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup与Requests提取HTML span标签文本的房产爬虫问题

房产爬虫文本提取&功能修正方案

现有代码核心问题

  • 搜索范围错误:你使用全局soup对象检索info-right、values等节点,每次都会返回页面第一个匹配的元素,无法对应到当前遍历的单个房源,需要将搜索范围限定在当前container容器内
  • CSS选择器用法错误:select_one方法如果传两个参数,第二个参数是命名空间,不是类筛选条件,要么直接写完整CSS选择器span.h-money,要么用find方法传class_参数
  • 未提取文本内容:你获取到DOM节点对象后没有调用.text属性,导致输出的是完整标签结构而不是内部文本

核心代码修正

from bs4 import BeautifulSoup
import requests
import mysql.connector

def main():
    
    list_price = []
    list_info_extra = []
    list_descrip = []
    list_url = []

    #Connection and cursor creation
    mydb = mysql.connector.connect(host="localhost", user="guilherme", passwd="fadel_gui", database="dawn18")
    cursor = mydb.cursor()
    if mydb.cursor:
        print("Connected to database")

    headers = ({'User-Agent': 'Mozilla/5.0 (Windows NT 6.1) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/41.0.2228 Safari/537.36'})

    URL = ["https://www.imobiliariapadreanchieta.com.br/imoveis/a-venda/apartamento/curitiba/bigorrilho", 
    "https://www.imobiliariapadreanchieta.com.br/imoveis/a-venda/apartamento/curitiba/bigorrilho?pagina=2", 
    "https://www.imobiliariapadreanchieta.com.br/imoveis/a-venda/apartamento/curitiba/bigorrilho?pagina=3"]

    for page_url in URL:
        
        response = requests.get(page_url, headers=headers)
        soup = BeautifulSoup(response.text, 'html.parser')
        house_containers = soup.find_all('div', class_= "col-sm-12 col-lg-6 box-align")

        if house_containers:
            for container in house_containers:
                # 提取房源价格
                price_node = container.find('div', class_="info-left")
                price = price_node.text.strip() if price_node else 'No info'

                # 提取Condomínio/IPTU信息
                info_right_node = container.select_one('div.info-right.text-xs-right p span.h-money')
                info_right = info_right_node.text.strip() if info_right_node else 'No info'

                # 提取其他费用信息
                info_containers = container.find_all('div', class_="values")
                info_apart_list = []
                for info in info_containers:
                    get_info = info.select_one('span.h-money')
                    if get_info:
                        info_apart_list.append(get_info.text.strip())
                    else:
                        info_apart_list.append('No info')

                # 提取房源链接,补充https前缀保证可直接访问
                url_node = container.find('a')
                url_imovel = 'https://www.imobiliariapadreanchieta.com.br' + url_node['href'] if url_node else 'No info'

                # 提取房间数、卫浴数、主卧数,对应站点类名分别为quartos、banheiros、suites
                bedroom_node = container.select_one('li.features__item--bedrooms span')
                bedroom = bedroom_node.text.strip() if bedroom_node else 'No info'
                bathroom_node = container.select_one('li.features__item--bathrooms span')
                bathroom = bathroom_node.text.strip() if bathroom_node else 'No info'
                suite_node = container.select_one('li.features__item--suites span')
                suite = suite_node.text.strip() if suite_node else 'No info'
                
                # 输出测试
                print(f"价格:{price}")
                print(f"Condomínio费用:{info_right}")
                print(f"房间数:{bedroom},卫浴数:{bathroom},主卧数:{suite}")
                print(f"房源链接:{url_imovel}")
                print("\n")

if __name__ == "__main__":
    main()

注意事项

  • 所有节点提取都加了非空判断,避免部分房源缺少对应字段时抛出异常
  • 提取内容加了.strip()方法,去掉多余的换行、空格字符
  • 你可以根据自己的需求把提取到的字段存入之前定义的列表或者直接写入数据库

内容的提问来源于stack exchange,提问作者Guilherme Celli Fadel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.30 06:54:03