如何解决Scraping导出CSV文件为空的问题
解决爬取fotocasa生成空casas.csv的问题
1. 缩进错误导致数据无法写入
你的for循环写在了with代码块外部,with语句执行完毕后文件会自动关闭,此时调用thewriter.writerow()根本无法向文件写入任何内容。
修正方式:将for循环缩进,放到with块内部,同时避免用list作为变量名(它是Python内置关键字):
with open('casas.csv', 'w', encoding='utf8', newline='') as f: thewriter = writer(f) header = ['Titulo', 'Precio', 'Metros', 'Telefono'] thewriter.writerow(header) # 将循环缩进至with内部 for list_item in lists: # 爬取逻辑... thewriter.writerow(info)
2. 未绕过网站反爬机制
fotocasa会检测请求的User-Agent标识,直接用requests.get()发送请求会被识别为爬虫,返回的页面不包含房源数据,导致lists = soup.find_all(...)得到空列表。
解决方法:添加请求头伪装成浏览器:
headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } page = requests.get(url, headers=headers)
3. 元素查找缺乏容错处理
如果页面中某个房源缺少特定元素(比如未显示电话),find()会返回None,直接调用.text会抛出AttributeError,导致程序中断,无法继续爬取后续数据。
优化方式:添加判断逻辑,处理元素不存在的情况:
titulo_elem = list_item.find('a', class_='re-CardPackAdvance-info-container') titulo = titulo_elem.text.replace('\n', '').strip() if titulo_elem else '无标题' precio_elem = list_item.find('span', class_='re-CardPrice') precio = precio_elem.text.replace('\n', '').strip() if precio_elem else '无价格' metros_elem = list_item.find('span', class_='re-CardFeaturesWithIcons-feature-icon--surface') metros = metros_elem.text.replace('\n', '').strip() if metros_elem else '无面积' telefono_elem = list_item.find('a', class_='re-CardContact-phone') telefono = telefono_elem.text.replace('\n', '').strip() if telefono_elem else '无电话'
完整修正后的代码
import requests from bs4 import BeautifulSoup from csv import writer url = 'https://www.fotocasa.es/es/alquiler/todas-las-casas/girona-provincia/todas-las-zonas/l' # 添加请求头绕过反爬 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } page = requests.get(url, headers=headers) soup = BeautifulSoup(page.content, 'html.parser') lists = soup.find_all('section', class_='re-CardPackAdvance') with open('casas.csv', 'w', encoding='utf8', newline='') as f: thewriter = writer(f) header = ['Titulo', 'Precio', 'Metros', 'Telefono'] thewriter.writerow(header) # 循环置于with内部,确保文件处于打开状态 for list_item in lists: # 容错处理每个元素 titulo_elem = list_item.find('a', class_='re-CardPackAdvance-info-container') titulo = titulo_elem.text.replace('\n', '').strip() if titulo_elem else '无标题' precio_elem = list_item.find('span', class_='re-CardPrice') precio = precio_elem.text.replace('\n', '').strip() if precio_elem else '无价格' metros_elem = list_item.find('span', class_='re-CardFeaturesWithIcons-feature-icon--surface') metros = metros_elem.text.replace('\n', '').strip() if metros_elem else '无面积' telefono_elem = list_item.find('a', class_='re-CardContact-phone') telefono = telefono_elem.text.replace('\n', '').strip() if telefono_elem else '无电话' info = [titulo, precio, metros, telefono] thewriter.writerow(info)
内容的提问来源于stack exchange,提问作者Hercules
相关产品推荐
相关产品推荐

