BeautifulSoup返回空列表,Python爬虫导出CSV仅含表头无数据求助
问题描述
我是Python网页爬虫初学者,运行下方代码时,BeautifulSoup获取的lists为空列表,终端无数据输出,生成的housing.csv文件仅写入表头,无其他数据,请求帮助解决该问题。
from bs4 import BeautifulSoup import requests from csv import writer url = "https://www.pararius.com/apartments/amsterdam" page = requests.get(url) soup = BeautifulSoup(page.content, 'html.parser') #za sve sekcije koje imaju class name listing-search-item vadim ih u list lists= soup.find_all('section', class_="listing-search-item") with open('housing.csv', 'w', encoding='utf8', newline='') as f: thewriter = writer(f) header = ['Title', 'Location', 'Price', 'Area'] thewriter.writerow(header) for list in lists: title = list.find('a', class_="listing-search-item__link listing-search-item__link--title").text.replace('\n', '') location = list.find('div', class_="entity-description-itemCaption").text.replace('\n', '') price = list.find('div', class_="listing-search-item__price").text.replace('\n', '') area = list.find('li', class_="illustrated-features__item--surface-area").text.replace('\n', '') info = [title, location, price, area] thewriter.writerow(info) print(lists)
问题分析与解决
核心原因
目标网站pararius.com存在基础反爬机制,直接使用requests.get()发送请求会被识别为非浏览器请求,返回的页面并非包含房源信息的真实内容,导致BeautifulSoup无法匹配到目标元素,最终lists为空。此外原代码中location对应的class名称存在拼写错误,也会导致元素查找失败。
解决步骤
- 添加请求头模拟浏览器:为
requests.get()添加User-Agent参数,模拟真实浏览器的请求标识,绕过基础反爬拦截。 - 修正class名称错误:将
location查找的class从entity-description-itemCaption改为真实页面中的entity-description-item__caption。 - 优化文本处理与异常处理:用
.strip()替代.replace('\n', '')更彻底清理文本,添加异常处理避免单个字段缺失导致程序中断。
修改后的代码
from bs4 import BeautifulSoup import requests from csv import writer url = "https://www.pararius.com/apartments/amsterdam" # 添加请求头,模拟Chrome浏览器(可根据自身浏览器调整User-Agent) headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } page = requests.get(url, headers=headers) soup = BeautifulSoup(page.content, 'html.parser') lists = soup.find_all('section', class_="listing-search-item") with open('housing.csv', 'w', encoding='utf8', newline='') as f: thewriter = writer(f) header = ['Title', 'Location', 'Price', 'Area'] thewriter.writerow(header) for item in lists: # 避免使用Python关键字list作为变量名 try: title = item.find('a', class_="listing-search-item__link listing-search-item__link--title").text.strip() location = item.find('div', class_="entity-description-item__caption").text.strip() price = item.find('div', class_="listing-search-item__price").text.strip() area = item.find('li', class_="illustrated-features__item--surface-area").text.strip() info = [title, location, price, area] thewriter.writerow(info) except AttributeError: # 跳过缺失字段的房源条目,避免程序崩溃 continue print(f"共抓取到{len(lists)}条房源数据")
额外说明
- 若后续再次出现拦截,可尝试添加更多请求头字段(如
Accept-Language),或使用time.sleep()添加请求间隔。 - 变量名避免使用Python内置关键字(如原代码中的
list),提升代码可读性与规范性。
内容的提问来源于stack exchange,提问作者Zoran Relić
相关产品推荐
相关产品推荐

