You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何修复Kununu多页数据提取的重复输出与死循环问题

修复kununu招聘数据爬取的重复与无限循环问题

嘿,我来帮你解决这个kununu数据爬取的问题!你的代码现在碰到两个头疼的问题:数据重复输出,还有程序停不下来无限运行。咱们一步步拆解问题,然后给出修复方案:

问题根源分析

  1. 无限循环:原代码靠判断new_comments是否为空来终止循环,但kununu的分页机制是——当你访问超过实际最大页码的页面时,它不会返回空页面,而是返回最后一页的内容。这样new_comments永远不为空,循环就停不下来。
  2. 数据重复:因为循环停不下来,会反复爬取最后一页的内容,导致data里堆满重复项;另外原代码直接抓取整个评论块的文本,没有拆分你需要的结构化字段,提取的内容也是杂乱无章的。

修复后的完整代码

import requests
from bs4 import BeautifulSoup
import re

def extract_application_data(soup):
    """提取单页内所有申请的结构化数据"""
    applications = []
    # 定位每个申请的容器(需根据实际页面结构调整class)
    application_containers = soup.find_all('div', class_='review-item')
    
    for container in application_containers:
        app_data = {}
        
        # 提取FIRMA:页面是大众的专区,默认值为Volkswagen,也可从页面元素提取
        firma_elem = container.find('span', class_='company-name')
        app_data['FIRMA'] = firma_elem.get_text(strip=True) if firma_elem else 'Volkswagen'
        
        # 提取STADT:定位城市相关元素
        stadt_elem = container.find('span', class_='location')
        app_data['STADT'] = stadt_elem.get_text(strip=True) if stadt_elem else None
        
        # 提取BEWORBEN FÜR POSITION:申请的职位
        position_elem = container.find('div', class_='job-title')
        app_data['BEWORBEN FÜR POSITION'] = position_elem.get_text(strip=True) if position_elem else None
        
        # 提取JAHR DER BEWERBUNG:从日期中提取年份
        date_elem = container.find('span', class_='review-date')
        date_text = date_elem.get_text(strip=True) if date_elem else ''
        year_match = re.search(r'\d{4}', date_text)
        app_data['JAHR DER BEWERBUNG'] = year_match.group() if year_match else None
        
        # 提取ERGEBNIS:申请结果(如angenommen/abgelehnt)
        result_elem = container.find('div', class_='application-outcome')
        app_data['ERGEBNIS'] = result_elem.get_text(strip=True) if result_elem else None
        
        applications.append(app_data)
    return applications

def main():
    all_data = []
    seen_app_ids = set()  # 用于去重的唯一标识集合
    base_url = 'https://www.kununu.com/de/volkswagen/bewerbung'
    max_pages = 50  # 最大页码限制,防止无限循环
    current_page = 1
    
    with requests.Session() as session:
        # 模拟浏览器请求头,避免被反爬拦截
        session.headers.update({
            'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
            'x-requested-with': 'XMLHttpRequest'
        })
        
        while current_page <= max_pages:
            print(f"正在处理第 {current_page} 页...")
            page_url = f"{base_url}/{current_page}"
            response = session.get(page_url)
            
            # 检查请求是否成功
            if response.status_code != 200:
                print(f"第 {current_page} 页请求失败,状态码: {response.status_code}")
                break
            
            soup = BeautifulSoup(response.text, 'html.parser')
            page_applications = extract_application_data(soup)
            
            # 如果当前页没有数据,终止循环
            if not page_applications:
                print(f"第 {current_page} 页无数据,停止爬取")
                break
            
            # 去重处理:过滤已爬取过的申请
            unique_apps = []
            for app in page_applications:
                # 用职位+城市+年份生成唯一标识
                app_id = f"{app['BEWORBEN FÜR POSITION']}_{app['STADT']}_{app['JAHR DER BEWERBUNG']}"
                if app_id not in seen_app_ids:
                    seen_app_ids.add(app_id)
                    unique_apps.append(app)
            
            # 如果当前页没有新数据,说明已经到最后一页
            if not unique_apps:
                print("没有新的申请数据,停止爬取")
                break
            
            # 添加去重后的数据
            all_data.extend(unique_apps)
            print(f"新增 {len(unique_apps)} 条有效数据,累计 {len(all_data)} 条")
            
            # 检查是否有下一页(通过分页按钮判断)
            next_page_btn = soup.find('a', class_='pagination-next')
            if not next_page_btn:
                print("未找到下一页按钮,停止爬取")
                break
            
            current_page += 1
    
    # 输出最终结果
    print("\n===== 提取完成 =====")
    for idx, app in enumerate(all_data, 1):
        print(f"\n申请 {idx}:")
        for key, value in app.items():
            print(f"- {key}: {value}")

if __name__ == "__main__":
    main()

关键改动说明

  1. 结构化字段提取:不再抓取整个文本块,而是针对每个目标字段(FIRMA、STADT等)单独解析对应的HTML元素,得到清晰的结构化数据。
  2. 多层循环终止机制:
    • 设置max_pages上限,兜底防止无限循环;
    • 检查请求状态码,失败则终止;
    • 检查当前页是否有数据,无数据则终止;
    • 检查是否有下一页按钮,没有则终止;
    • 检查当前页是否有新数据(去重后为空),则终止。
  3. 数据去重:用seen_app_ids集合存储每条申请的唯一标识(职位+城市+年份),避免重复添加相同的申请。
  4. 反爬优化:添加User-Agent模拟浏览器请求,降低被网站拦截的概率。

注意事项

  • HTML class调整:代码中的class(如review-item、job-title)是基于常见页面结构假设的,你需要打开浏览器开发者工具,查看kununu实际的HTML元素class,然后修改代码中对应的参数。
  • 请求频率控制:可以在每页请求后添加time.sleep(1)(需导入time库),降低爬取频率,避免被封禁IP。
  • 遵守网站规则:确保你的爬取行为符合kununu的robots.txt规则,不要过度爬取。

内容的提问来源于stack exchange,提问作者codecodecode

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 09:10:35