Python爬取Airbnb评论数据异常:代码直接进入except分支求助
解决Airbnb爬虫代码无法进入try块的问题
Hey there, let's figure out why your Airbnb scraper keeps hitting the except blocks instead of pulling the data you want. I've gone through your code and spotted a couple key issues that are causing this:
1. 反爬拦截导致页面内容异常
Airbnb有很强的反爬机制,当你直接用requests.get()发送请求而不带任何请求头时,服务器很容易识别出你是爬虫,返回的不是真实的房源页面内容——这就导致你的soup对象里根本找不到你要定位的元素,所有try块自然都会触发异常。
修复方案:添加模拟浏览器的请求头
给请求加上浏览器标识、语言偏好等头信息,让请求看起来更像真实用户访问:
def get_page(url): # 模拟Chrome浏览器的请求头 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Accept-Language': 'en-US,en;q=0.9', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8' } response = requests.get(url, headers=headers) if not response.ok: print('Server responded:', response.status_code) return None # 返回None提示页面获取失败 else: soup = BeautifulSoup(response.text, 'html.parser') return soup
2. 错误使用find_all定位单个元素
你的代码里用find_all()来定位房源名称、总评论数这类唯一元素,但find_all()返回的是元素列表(ResultSet),直接对列表调用.text会触发AttributeError,直接跳转到except块。
修复方案:用find()定位单个元素,用find_all()遍历批量元素
- 对于房源标题、总评论数这种单个元素,改用
soup.find(); - 对于评论内容、评论者姓名这类批量元素,用
find_all()获取列表后再遍历处理。
3. 冗余代码与数据收集缺失
你的代码里有重复的try-except块(比如重复获取total_reviews和comment_date),而且get_detail_data函数没有返回或保存收集到的数据。下面是优化后的完整函数:
def get_detail_data(soup): if not soup: print("Failed to get page content") return # 提取房源名称 try: title = soup.find('span', class_="_18hrqvin").text.strip() except AttributeError: title = 'empty' print(f"Title: {title}") # 提取总评论数 try: total_reviews = soup.find('span', class_="_krjbj").text.strip() except AttributeError: total_reviews = 'empty total reviews' print(f"Total Reviews: {total_reviews}") # 批量提取评论相关数据 comments_data = [] comment_containers = soup.find_all('div', class_="_czm8crp") # 评论内容容器 commenter_names = soup.find_all('div', class_="_1p3joamp") # 评论者姓名 comment_dates = soup.find_all('span', class_="_1jlnvra2") # 评论日期 for idx, container in enumerate(comment_containers): # 提取单条评论 try: comment = container.text.strip() except AttributeError: comment = 'empty comment' # 提取对应评论者姓名(避免索引越界) try: commenter = commenter_names[idx].text.strip() if idx < len(commenter_names) else 'empty name' except AttributeError: commenter = 'empty name' # 提取对应评论日期 try: date = comment_dates[idx].text.strip() if idx < len(comment_dates) else 'empty date' except AttributeError: date = 'empty date' comments_data.append({ 'title': title, 'total_reviews': total_reviews, 'commenter_name': commenter, 'comment_date': date, 'comment': comment }) # 将数据保存到CSV文件 with open('airbnb_reviews.csv', 'w', newline='', encoding='utf-8') as f: writer = csv.DictWriter(f, fieldnames=['title', 'total_reviews', 'commenter_name', 'comment_date', 'comment']) writer.writeheader() writer.writerows(comments_data) print(f"Successfully saved {len(comments_data)} comments to CSV")
额外注意事项
- Airbnb反爬升级:即使加了请求头,频繁请求也可能被封IP。可以考虑添加请求延迟、使用代理IP,或者用Selenium模拟真实浏览器(Airbnb很多内容是JS渲染的,静态请求可能拿不到完整数据);
- 类名变动:Airbnb经常会修改页面元素的类名(比如
_18hrqvin这种随机类名),如果代码突然失效,记得检查页面源码更新选择器; - 合规性:请确保你的爬虫行为符合Airbnb的服务条款,未经授权的爬虫可能违反平台规定。
内容的提问来源于stack exchange,提问作者M.Akram
相关产品推荐
相关产品推荐

