You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas Concat拼接数据失败:DataFrame始终为空的问题求助

问题分析与解决方案

你的核心问题出在DataFrame拼接的调用方式错误,同时缺少结果赋值操作,导致数据无法写入DataFrame。以下是具体问题点和修正方案:

问题点拆解

  1. 错误调用df.concat:concat是pandas的顶层函数,正确写法是pd.concat(),而非DataFrame对象的方法
  2. 未赋值拼接结果:concat默认返回新的DataFrame,不会修改原对象,必须将结果重新赋值给df
  3. 直接传入字典无效:需要将单条数据转换为DataFrame/Series后再拼接
  4. 初始DataFrame带空行:会导致最终结果保留多余的空数据

修正方案1:修正concat用法

import requests
from bs4 import BeautifulSoup
import pandas as pd

url = 'https://www.airbnb.com/s/Honolulu--HI--United-States/homes?tab_id=home_tab&refinement_paths%5B%5D=%2Fhomes&flexible_trip_lengths%5B%5D=one_week&price_filter_input_type=0&query=Honolulu%2C%20HI&place_id=ChIJTUbDjDsYAHwRbJen81_1KEs&date_picker_type=calendar&checkin=2022-10-08&checkout=2022-10-09&source=structured_search_input_header&search_type=autocomplete_click'
page = requests.get(url, headers={'User-agent': 'your bot 0.1'})
soup = BeautifulSoup(page.text, 'lxml')

# 初始化空列的DataFrame,避免多余空行
df = pd.DataFrame(columns=['Links', 'Title', 'Price', 'Rating'])
postings = soup.findAll('div', class_='c4mnd7m dir dir-ltr')

for post in postings:
    try:
        title = post.find('div', class_='t1jojoys dir dir-ltr').text
        link = 'https://www.airbnb.com/' + post.find('a', class_='ln2bl2p dir dir-ltr').get('href')
        price = post.find('span', class_='a8jt5op dir dir-ltr').text
        rating = post.find('span', class_='ru0q88m dir dir-ltr').text
        
        # 将单条数据转为DataFrame,再与原df拼接并赋值
        new_row = pd.DataFrame({'Links': [link], 'Title': [title], 'Price': [price], 'Rating': [rating]})
        df = pd.concat([df, new_row], ignore_index=True)
    except Exception as e:
        # 打印异常方便排查问题
        print(f"处理房源出错: {e}")
        pass

print(df)

修正方案2:列表收集法(推荐,效率更高)

循环中频繁拼接DataFrame效率较低,更推荐先将所有数据存入列表,最后一次性转换为DataFrame:

import requests
from bs4 import BeautifulSoup
import pandas as pd

url = 'https://www.airbnb.com/s/Honolulu--HI--United-States/homes?tab_id=home_tab&refinement_paths%5B%5D=%2Fhomes&flexible_trip_lengths%5B%5D=one_week&price_filter_input_type=0&query=Honolulu%2C%20HI&place_id=ChIJTUbDjDsYAHwRbJen81_1KEs&date_picker_type=calendar&checkin=2022-10-08&checkout=2022-10-09&source=structured_search_input_header&search_type=autocomplete_click'
page = requests.get(url, headers={'User-agent': 'your bot 0.1'})
soup = BeautifulSoup(page.text, 'lxml')

# 初始化列表存储所有房源数据
data_list = []
postings = soup.findAll('div', class_='c4mnd7m dir dir-ltr')

for post in postings:
    try:
        title = post.find('div', class_='t1jojoys dir dir-ltr').text
        link = 'https://www.airbnb.com/' + post.find('a', class_='ln2bl2p dir dir-ltr').get('href')
        price = post.find('span', class_='a8jt5op dir dir-ltr').text
        rating = post.find('span', class_='ru0q88m dir dir-ltr').text
        
        # 将单条数据加入列表
        data_list.append({
            'Links': link,
            'Title': title,
            'Price': price,
            'Rating': rating
        })
    except Exception as e:
        print(f"处理房源出错: {e}")
        pass

# 最后一次性转换为DataFrame
df = pd.DataFrame(data_list)
print(df)

内容的提问来源于stack exchange,提问作者Japo Japic

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 23:20:26