You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解决Scraping导出CSV文件为空的问题

解决爬取fotocasa生成空casas.csv的问题

1. 缩进错误导致数据无法写入

你的for循环写在了with代码块外部,with语句执行完毕后文件会自动关闭,此时调用thewriter.writerow()根本无法向文件写入任何内容。

修正方式:将for循环缩进,放到with块内部,同时避免用list作为变量名(它是Python内置关键字):

with open('casas.csv', 'w', encoding='utf8', newline='') as f:
    thewriter = writer(f)
    header = ['Titulo', 'Precio', 'Metros', 'Telefono']
    thewriter.writerow(header)
    # 将循环缩进至with内部
    for list_item in lists:
        # 爬取逻辑...
        thewriter.writerow(info)

2. 未绕过网站反爬机制

fotocasa会检测请求的User-Agent标识,直接用requests.get()发送请求会被识别为爬虫,返回的页面不包含房源数据,导致lists = soup.find_all(...)得到空列表。

解决方法:添加请求头伪装成浏览器:

headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
}
page = requests.get(url, headers=headers)

3. 元素查找缺乏容错处理

如果页面中某个房源缺少特定元素(比如未显示电话),find()会返回None,直接调用.text会抛出AttributeError,导致程序中断,无法继续爬取后续数据。

优化方式:添加判断逻辑,处理元素不存在的情况:

titulo_elem = list_item.find('a', class_='re-CardPackAdvance-info-container')
titulo = titulo_elem.text.replace('\n', '').strip() if titulo_elem else '无标题'

precio_elem = list_item.find('span', class_='re-CardPrice')
precio = precio_elem.text.replace('\n', '').strip() if precio_elem else '无价格'

metros_elem = list_item.find('span', class_='re-CardFeaturesWithIcons-feature-icon--surface')
metros = metros_elem.text.replace('\n', '').strip() if metros_elem else '无面积'

telefono_elem = list_item.find('a', class_='re-CardContact-phone')
telefono = telefono_elem.text.replace('\n', '').strip() if telefono_elem else '无电话'

完整修正后的代码

import requests
from bs4 import BeautifulSoup
from csv import writer

url = 'https://www.fotocasa.es/es/alquiler/todas-las-casas/girona-provincia/todas-las-zonas/l'
# 添加请求头绕过反爬
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
}
page = requests.get(url, headers=headers)
soup = BeautifulSoup(page.content, 'html.parser')
lists = soup.find_all('section', class_='re-CardPackAdvance')

with open('casas.csv', 'w', encoding='utf8', newline='') as f:
    thewriter = writer(f)
    header = ['Titulo', 'Precio', 'Metros', 'Telefono']
    thewriter.writerow(header)
    
    # 循环置于with内部,确保文件处于打开状态
    for list_item in lists:
        # 容错处理每个元素
        titulo_elem = list_item.find('a', class_='re-CardPackAdvance-info-container')
        titulo = titulo_elem.text.replace('\n', '').strip() if titulo_elem else '无标题'
        
        precio_elem = list_item.find('span', class_='re-CardPrice')
        precio = precio_elem.text.replace('\n', '').strip() if precio_elem else '无价格'
        
        metros_elem = list_item.find('span', class_='re-CardFeaturesWithIcons-feature-icon--surface')
        metros = metros_elem.text.replace('\n', '').strip() if metros_elem else '无面积'
        
        telefono_elem = list_item.find('a', class_='re-CardContact-phone')
        telefono = telefono_elem.text.replace('\n', '').strip() if telefono_elem else '无电话'
        
        info = [titulo, precio, metros, telefono]
        thewriter.writerow(info)

内容的提问来源于stack exchange,提问作者Hercules

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 01:40:32