You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

BeautifulSoup返回空列表,Python爬虫导出CSV仅含表头无数据求助

问题描述

我是Python网页爬虫初学者,运行下方代码时,BeautifulSoup获取的lists为空列表,终端无数据输出,生成的housing.csv文件仅写入表头,无其他数据,请求帮助解决该问题。

from bs4 import BeautifulSoup
import requests
from csv import writer

url = "https://www.pararius.com/apartments/amsterdam"

page = requests.get(url)

soup = BeautifulSoup(page.content, 'html.parser')

#za sve sekcije koje imaju class name listing-search-item vadim ih u list
lists= soup.find_all('section', class_="listing-search-item")

with open('housing.csv', 'w', encoding='utf8', newline='') as f:
    thewriter = writer(f)
    
    header = ['Title', 'Location', 'Price', 'Area']
    thewriter.writerow(header)
    for list in lists:

        title = list.find('a', class_="listing-search-item__link listing-search-item__link--title").text.replace('\n', '')
        location = list.find('div', class_="entity-description-itemCaption").text.replace('\n', '')
        price = list.find('div', class_="listing-search-item__price").text.replace('\n', '')
        area = list.find('li', class_="illustrated-features__item--surface-area").text.replace('\n', '')
        info = [title, location, price, area]
        thewriter.writerow(info)

print(lists)
问题分析与解决

核心原因

目标网站pararius.com存在基础反爬机制,直接使用requests.get()发送请求会被识别为非浏览器请求,返回的页面并非包含房源信息的真实内容,导致BeautifulSoup无法匹配到目标元素,最终lists为空。此外原代码中location对应的class名称存在拼写错误,也会导致元素查找失败。

解决步骤

  1. 添加请求头模拟浏览器:为requests.get()添加User-Agent参数,模拟真实浏览器的请求标识,绕过基础反爬拦截。
  2. 修正class名称错误:将location查找的class从entity-description-itemCaption改为真实页面中的entity-description-item__caption。
  3. 优化文本处理与异常处理:用.strip()替代.replace('\n', '')更彻底清理文本,添加异常处理避免单个字段缺失导致程序中断。

修改后的代码

from bs4 import BeautifulSoup
import requests
from csv import writer

url = "https://www.pararius.com/apartments/amsterdam"

# 添加请求头,模拟Chrome浏览器(可根据自身浏览器调整User-Agent)
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
}

page = requests.get(url, headers=headers)

soup = BeautifulSoup(page.content, 'html.parser')

lists = soup.find_all('section', class_="listing-search-item")

with open('housing.csv', 'w', encoding='utf8', newline='') as f:
    thewriter = writer(f)
    
    header = ['Title', 'Location', 'Price', 'Area']
    thewriter.writerow(header)
    for item in lists:  # 避免使用Python关键字list作为变量名
        try:
            title = item.find('a', class_="listing-search-item__link listing-search-item__link--title").text.strip()
            location = item.find('div', class_="entity-description-item__caption").text.strip()
            price = item.find('div', class_="listing-search-item__price").text.strip()
            area = item.find('li', class_="illustrated-features__item--surface-area").text.strip()
            info = [title, location, price, area]
            thewriter.writerow(info)
        except AttributeError:
            # 跳过缺失字段的房源条目,避免程序崩溃
            continue

print(f"共抓取到{len(lists)}条房源数据")

额外说明

  • 若后续再次出现拦截,可尝试添加更多请求头字段(如Accept-Language),或使用time.sleep()添加请求间隔。
  • 变量名避免使用Python内置关键字(如原代码中的list),提升代码可读性与规范性。

内容的提问来源于stack exchange,提问作者Zoran Relić

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 17:10:20