使用concat和_append时遇DataFrame无append属性错误且输出空集如何解决?
问题修复方案
一、解决AttributeError: 'DataFrame' object has no attribute 'append'
Pandas 2.0及以上版本已移除DataFrame.append()方法,如果你之前代码中使用了final.append(dataset),直接替换为pd.concat()即可,就像你当前代码里的写法:
final = pd.concat([final, dataset], ignore_index=True)
确保代码中不再调用append()方法,该错误即可解决。
二、解决数据集为空的问题
你的代码存在三个核心错误,导致无法爬取到有效数据:
1. 循环缩进错误
for j in range(1,800):语句后未添加缩进,导致后续的请求、解析代码都不在循环体内,只会执行一次且无法遍历页码。需要将循环内的所有代码(从headers定义到concat操作)全部缩进,纳入循环块中。
2. URL未格式化页码
请求URL中的{}没有替换为当前循环的页码j,导致请求的是无效链接。需用字符串格式化替换页码:
response = requests.get(f'https://www.ambitionbox.com/list-of-companies?campaign=desktop_nav&page={j}', headers=headers).text
或使用.format(j)写法:
response = requests.get('https://www.ambitionbox.com/list-of-companies?campaign=desktop_nav&page={}'.format(j), headers=headers).text
3. User-Agent格式错误
你的User-Agent存在拼写错误和非法空格(如Apple WeKit应为AppleWebKit,含中文空格),修正为标准格式:
headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 6.3; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/80.0.3987.162 Safari/537.36'}
可选优化:添加异常处理
为避免单个页面请求失败导致程序中断,可添加异常捕获逻辑:
try: response = requests.get(..., timeout=10) response.raise_for_status() # 检查请求是否成功 # 解析代码... except Exception as e: print(f"第{j}页爬取失败: {e}") continue
修正后的完整代码
import pandas as pd import requests from bs4 import BeautifulSoup import time final = pd.DataFrame() headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 6.3; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/80.0.3987.162 Safari/537.36'} for j in range(1, 800): try: # 格式化URL并发起请求 url = f'https://www.ambitionbox.com/list-of-companies?campaign=desktop_nav&page={j}' response = requests.get(url, headers=headers, timeout=10) response.raise_for_status() soup = BeautifulSoup(response.text, 'html.parser') company = soup.find_all('div', class_='company-content-wrapper') # 每次循环初始化列表,避免数据累积 name = [] rating = [] reviews = [] ctype = [] hq = [] how_old = [] no_of_employee = [] for i in company: # 提取名称,增加空值判断 name_tag = i.find('h2') name.append(name_tag.text.strip() if name_tag else "N/A") # 提取评分 rating_tag = i.find('p', class_='rating') rating.append(rating_tag.text.strip() if rating_tag else "N/A") # 提取评论数 review_tag = i.find('a', class_='review-count') reviews.append(review_tag.text.strip() if review_tag else "N/A") info_entities = i.find_all('p', class_='infoEntity') ctype.append(info_entities[0].text.strip() if len(info_entities) > 0 else "N/A") hq.append(info_entities[1].text.strip() if len(info_entities) > 1 else "N/A") how_old.append(info_entities[2].text.strip() if len(info_entities) > 2 else "N/A") no_of_employee.append(info_entities[3].text.strip() if len(info_entities) > 3 else "N/A") dataset = pd.DataFrame({ 'Name': name, 'Rating': rating, 'reviews': reviews, 'Company Type': ctype, 'Headquaters': hq, 'Company Age': how_old, 'No. of Employee': no_of_employee }) final = pd.concat([final, dataset], ignore_index=True) print(f"第{j}页爬取完成,当前总数据量:{len(final)}") time.sleep(1) # 添加请求间隔,避免反爬 except Exception as e: print(f"第{j}页爬取失败: {str(e)}") continue # 保存数据,去除索引列 final.to_csv('Company_Dataset.csv', index=False) print("数据保存完成!")
额外注意事项
- AmbitionBox存在反爬机制,频繁请求会触发IP封禁,建议保留
time.sleep(1)的请求间隔。 - 若后续网站更新页面结构,需同步调整BeautifulSoup的查找参数(如class名称)。
内容的提问来源于stack exchange,提问作者Milan Kumawat
相关产品推荐
相关产品推荐

