You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬取Airbnb评论数据异常:代码直接进入except分支求助

解决Airbnb爬虫代码无法进入try块的问题

Hey there, let's figure out why your Airbnb scraper keeps hitting the except blocks instead of pulling the data you want. I've gone through your code and spotted a couple key issues that are causing this:

1. 反爬拦截导致页面内容异常

Airbnb有很强的反爬机制,当你直接用requests.get()发送请求而不带任何请求头时,服务器很容易识别出你是爬虫,返回的不是真实的房源页面内容——这就导致你的soup对象里根本找不到你要定位的元素,所有try块自然都会触发异常。

修复方案:添加模拟浏览器的请求头
给请求加上浏览器标识、语言偏好等头信息,让请求看起来更像真实用户访问:

def get_page(url):
    # 模拟Chrome浏览器的请求头
    headers = {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
        'Accept-Language': 'en-US,en;q=0.9',
        'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8'
    }
    response = requests.get(url, headers=headers)
    if not response.ok:
        print('Server responded:', response.status_code)
        return None  # 返回None提示页面获取失败
    else:
        soup = BeautifulSoup(response.text, 'html.parser')
        return soup

2. 错误使用find_all定位单个元素

你的代码里用find_all()来定位房源名称、总评论数这类唯一元素,但find_all()返回的是元素列表(ResultSet),直接对列表调用.text会触发AttributeError,直接跳转到except块。

修复方案:用find()定位单个元素,用find_all()遍历批量元素

  • 对于房源标题、总评论数这种单个元素,改用soup.find();
  • 对于评论内容、评论者姓名这类批量元素,用find_all()获取列表后再遍历处理。

3. 冗余代码与数据收集缺失

你的代码里有重复的try-except块(比如重复获取total_reviews和comment_date),而且get_detail_data函数没有返回或保存收集到的数据。下面是优化后的完整函数:

def get_detail_data(soup):
    if not soup:
        print("Failed to get page content")
        return
    
    # 提取房源名称
    try:
        title = soup.find('span', class_="_18hrqvin").text.strip()
    except AttributeError:
        title = 'empty'
        print(f"Title: {title}")
    
    # 提取总评论数
    try:
        total_reviews = soup.find('span', class_="_krjbj").text.strip()
    except AttributeError:
        total_reviews = 'empty total reviews'
        print(f"Total Reviews: {total_reviews}")
    
    # 批量提取评论相关数据
    comments_data = []
    comment_containers = soup.find_all('div', class_="_czm8crp")  # 评论内容容器
    commenter_names = soup.find_all('div', class_="_1p3joamp")   # 评论者姓名
    comment_dates = soup.find_all('span', class_="_1jlnvra2")    # 评论日期
    
    for idx, container in enumerate(comment_containers):
        # 提取单条评论
        try:
            comment = container.text.strip()
        except AttributeError:
            comment = 'empty comment'
        
        # 提取对应评论者姓名(避免索引越界)
        try:
            commenter = commenter_names[idx].text.strip() if idx < len(commenter_names) else 'empty name'
        except AttributeError:
            commenter = 'empty name'
        
        # 提取对应评论日期
        try:
            date = comment_dates[idx].text.strip() if idx < len(comment_dates) else 'empty date'
        except AttributeError:
            date = 'empty date'
        
        comments_data.append({
            'title': title,
            'total_reviews': total_reviews,
            'commenter_name': commenter,
            'comment_date': date,
            'comment': comment
        })
    
    # 将数据保存到CSV文件
    with open('airbnb_reviews.csv', 'w', newline='', encoding='utf-8') as f:
        writer = csv.DictWriter(f, fieldnames=['title', 'total_reviews', 'commenter_name', 'comment_date', 'comment'])
        writer.writeheader()
        writer.writerows(comments_data)
    
    print(f"Successfully saved {len(comments_data)} comments to CSV")

额外注意事项

  • Airbnb反爬升级:即使加了请求头,频繁请求也可能被封IP。可以考虑添加请求延迟、使用代理IP,或者用Selenium模拟真实浏览器(Airbnb很多内容是JS渲染的,静态请求可能拿不到完整数据);
  • 类名变动:Airbnb经常会修改页面元素的类名(比如_18hrqvin这种随机类名),如果代码突然失效,记得检查页面源码更新选择器;
  • 合规性:请确保你的爬虫行为符合Airbnb的服务条款,未经授权的爬虫可能违反平台规定。

内容的提问来源于stack exchange,提问作者M.Akram

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 23:52:41