使用Scrapfly爬取Indeed/Glassdoor评论遇错误及空DataFrame求助
解决Scrapfly爬取Indeed评论的空DataFrame及调用错误问题
1. 修复Scrapfly SDK调用错误
你遇到的'dict' object has no attribute 'method'错误是因为Scrapfly Python SDK调用方式不正确,新版本SDK要求使用ScrapeRequest对象作为参数,而非直接传入字典。
修改请求代码:
# 先导入ScrapeRequest from scrapfly.scrape import ScrapeRequest # 替换原来的response请求代码 response = client.scrape(ScrapeRequest(url=url, render_js=True))
注意:Indeed的评论内容是动态渲染的,必须开启render_js=True才能获取到完整数据。
2. 修正页面选择器(适配Indeed当前页面结构)
你的原始选择器已失效,Indeed当前的评论容器类名和属性有更新,调整选择器如下:
- 单个评论容器:
div.cmp-review-container - 标题:
h3.cmp-review-title - 作者与身份:
span.cmp-reviewer-job-title(格式如"Former Employee - Software Engineer") - 评分:
div.cmp-review-rating span的aria-label - Pros/Cons:
div.cmp-review-pros-cons下的对应子元素
3. 完整修正后的代码
!pip install beautifulsoup4 pandas lxml requests !git clone https://github.com/scrapfly/python-sdk.git %cd python-sdk !pip install . import os import pandas as pd from scrapfly.client import ScrapflyClient from scrapfly.scrape import ScrapeRequest from bs4 import BeautifulSoup # 设置Scrapfly API密钥 os.environ["SCRAPFLY_KEY"] = "your_key_here" SCRAPFLY_KEY = os.getenv("SCRAPFLY_KEY") client = ScrapflyClient(key=SCRAPFLY_KEY) # 初始化结果存储列表 lst = [] # 遍历页面采集数据 for i in range(0, 240, 20): print(f"Scraping page {i // 20 + 1}...") url = f"https://www.indeed.com/cmp/Airbnb/reviews?start={i}" try: # 使用正确的ScrapeRequest对象调用 response = client.scrape(ScrapeRequest(url=url, render_js=True)) soup = BeautifulSoup(response.content, 'lxml') # 定位单个评论容器 main_data = soup.find_all("div", class_="cmp-review-container") for data in main_data: # 提取标题 title = data.find("h3", class_="cmp-review-title").get_text(strip=True) if data.find("h3", class_="cmp-review-title") else None # 提取作者与身份 author_status = data.find("span", class_="cmp-reviewer-job-title").get_text(strip=True) if data.find("span", class_="cmp-reviewer-job-title") else "" if "-" in author_status: status, author = author_status.split("-", 1) status = status.strip() author = author.strip() else: status, author = None, None # 提取评论正文 review = data.find("div", class_="cmp-review-text").get_text(strip=True) if data.find("div", class_="cmp-review-text") else None # 提取Pros和Cons pros_cons = data.find("div", class_="cmp-review-pros-cons") pros = pros_cons.find("div", class_="cmp-review-pro-text").get_text(strip=True) if pros_cons and pros_cons.find("div", class_="cmp-review-pro-text") else None cons = pros_cons.find("div", class_="cmp-review-con-text").get_text(strip=True) if pros_cons and pros_cons.find("div", class_="cmp-review-con-text") else None # 提取评分 rating_elem = data.find("div", class_="cmp-review-rating") rating = rating_elem.find("span")["aria-label"].split(" ")[0] if rating_elem else None lst.append([title, author, status, pros, cons, review, rating]) except Exception as e: print(f"Error on page {i // 20 + 1}: {e}") # 创建DataFrame并保存 df = pd.DataFrame(lst, columns=['Title', 'Author', 'Status', 'Pros', 'Cons', 'Review', 'Rating']) print(df.head()) df.to_csv("/content/drive/MyDrive/indeed_reviews_scrapfly.csv", index=False)
4. 额外注意事项
- Glassdoor爬取逻辑类似:同样需要开启JS渲染,且要适配Glassdoor的页面选择器(比如评论容器用
div.review)。 - 请求频率控制:Scrapfly默认会处理反爬,但不要过度频繁请求,避免触发目标网站的额外限制。
- 选择器验证:如果再次出现空数据,建议手动打开目标页面,用浏览器开发者工具检查最新的HTML结构,更新选择器。
内容的提问来源于stack exchange,提问作者Soyeon Lee
相关产品推荐
相关产品推荐

