You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用concat和_append时遇DataFrame无append属性错误且输出空集如何解决?

问题修复方案

一、解决AttributeError: 'DataFrame' object has no attribute 'append'

Pandas 2.0及以上版本已移除DataFrame.append()方法,如果你之前代码中使用了final.append(dataset),直接替换为pd.concat()即可,就像你当前代码里的写法:

final = pd.concat([final, dataset], ignore_index=True)

确保代码中不再调用append()方法,该错误即可解决。

二、解决数据集为空的问题

你的代码存在三个核心错误,导致无法爬取到有效数据:

1. 循环缩进错误

for j in range(1,800):语句后未添加缩进,导致后续的请求、解析代码都不在循环体内,只会执行一次且无法遍历页码。需要将循环内的所有代码(从headers定义到concat操作)全部缩进,纳入循环块中。

2. URL未格式化页码

请求URL中的{}没有替换为当前循环的页码j,导致请求的是无效链接。需用字符串格式化替换页码:

response = requests.get(f'https://www.ambitionbox.com/list-of-companies?campaign=desktop_nav&page={j}', headers=headers).text

或使用.format(j)写法:

response = requests.get('https://www.ambitionbox.com/list-of-companies?campaign=desktop_nav&page={}'.format(j), headers=headers).text

3. User-Agent格式错误

你的User-Agent存在拼写错误和非法空格(如Apple WeKit应为AppleWebKit,含中文空格),修正为标准格式:

headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 6.3; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/80.0.3987.162 Safari/537.36'}

可选优化:添加异常处理

为避免单个页面请求失败导致程序中断,可添加异常捕获逻辑:

try:
    response = requests.get(..., timeout=10)
    response.raise_for_status()  # 检查请求是否成功
    # 解析代码...
except Exception as e:
    print(f"第{j}页爬取失败: {e}")
    continue

修正后的完整代码

import pandas as pd
import requests
from bs4 import BeautifulSoup
import time

final = pd.DataFrame()
headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 6.3; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/80.0.3987.162 Safari/537.36'}

for j in range(1, 800):
    try:
        # 格式化URL并发起请求
        url = f'https://www.ambitionbox.com/list-of-companies?campaign=desktop_nav&page={j}'
        response = requests.get(url, headers=headers, timeout=10)
        response.raise_for_status()
        soup = BeautifulSoup(response.text, 'html.parser')
        company = soup.find_all('div', class_='company-content-wrapper')
        
        # 每次循环初始化列表,避免数据累积
        name = []
        rating = []
        reviews = []
        ctype = []
        hq = []
        how_old = []
        no_of_employee = []
        
        for i in company:
            # 提取名称,增加空值判断
            name_tag = i.find('h2')
            name.append(name_tag.text.strip() if name_tag else "N/A")
            
            # 提取评分
            rating_tag = i.find('p', class_='rating')
            rating.append(rating_tag.text.strip() if rating_tag else "N/A")
            
            # 提取评论数
            review_tag = i.find('a', class_='review-count')
            reviews.append(review_tag.text.strip() if review_tag else "N/A")
            
            info_entities = i.find_all('p', class_='infoEntity')
            ctype.append(info_entities[0].text.strip() if len(info_entities) > 0 else "N/A")
            hq.append(info_entities[1].text.strip() if len(info_entities) > 1 else "N/A")
            how_old.append(info_entities[2].text.strip() if len(info_entities) > 2 else "N/A")
            no_of_employee.append(info_entities[3].text.strip() if len(info_entities) > 3 else "N/A")
        
        dataset = pd.DataFrame({
            'Name': name,
            'Rating': rating,
            'reviews': reviews,
            'Company Type': ctype,
            'Headquaters': hq,
            'Company Age': how_old,
            'No. of Employee': no_of_employee
        })
        
        final = pd.concat([final, dataset], ignore_index=True)
        print(f"第{j}页爬取完成,当前总数据量:{len(final)}")
        time.sleep(1)  # 添加请求间隔,避免反爬
        
    except Exception as e:
        print(f"第{j}页爬取失败: {str(e)}")
        continue

# 保存数据,去除索引列
final.to_csv('Company_Dataset.csv', index=False)
print("数据保存完成!")

额外注意事项

  • AmbitionBox存在反爬机制,频繁请求会触发IP封禁,建议保留time.sleep(1)的请求间隔。
  • 若后续网站更新页面结构,需同步调整BeautifulSoup的查找参数(如class名称)。

内容的提问来源于stack exchange,提问作者Milan Kumawat

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 15:12:33