Flask招聘爬虫500错误:DataFrame数组长度不一致问题排查
问题定位
ValueError: All arrays must be of the same length 本质是你爬取LinkedIn职位数据时,存储各字段(如职位标题、公司名、工作地点等)的列表长度不一致。比如爬了10条职位,但其中2条没抓到地点,地点列表长度为8,其他字段列表长度为10,pandas构建DataFrame时要求所有列长度必须相等,因此抛出错误。
修复方案
1. 字段提取时添加默认值
对每个字段的提取逻辑做容错处理,提取失败时填充默认值(如"N/A"),确保每个列表的长度和总职位数一致:
titles = [] companies = [] locations = [] job_cards = # 你的职位卡片选择器结果 for job in job_cards: # 提取职位标题,无则填N/A title = job.find("h3").text.strip() if job.find("h3") else "N/A" titles.append(title) # 提取公司名称,无则填N/A(注意LinkedIn的class可能随页面更新变化,需核对) company_elem = job.find("span", class_="company-name") company = company_elem.text.strip() if company_elem else "N/A" companies.append(company) # 提取工作地点,无则填N/A location_elem = job.find("span", class_="job-search-card__location") location = location_elem.text.strip() if location_elem else "N/A" locations.append(location)
2. 构建DataFrame前校验长度
在生成DataFrame前,强制统一所有列表的长度,避免遗漏导致的错误:
import pandas as pd # 获取所有字段列表的最大长度 max_length = max(len(titles), len(companies), len(locations)) # 补全较短的列表,用N/A填充 titles += ["N/A"] * (max_length - len(titles)) companies += ["N/A"] * (max_length - len(companies)) locations += ["N/A"] * (max_length - len(locations)) # 构建DataFrame df = pd.DataFrame({ "职位标题": titles, "公司名称": companies, "工作地点": locations })
3. 添加异常捕获调试
给字段提取环节加try-except,方便定位具体哪个职位的字段抓取失败:
for idx, job in enumerate(job_cards): try: title = job.find("h3").text.strip() except AttributeError: title = "N/A" print(f"第{idx+1}条职位:标题提取失败") titles.append(title) # 其他字段同理添加异常捕获
后续功能实现建议(职位/地点用户输入)
1. 修改Flask路由支持表单提交
from flask import Flask, render_template, request import urllib.parse app = Flask(__name__) @app.route("/", methods=["GET", "POST"]) def index(): if request.method == "POST": # 获取用户输入的关键词和地点 keyword = request.form.get("keyword").strip() location = request.form.get("location").strip() # 对参数做URL编码,避免特殊字符报错 encoded_keyword = urllib.parse.quote(keyword) encoded_location = urllib.parse.quote(location) # 构建LinkedIn搜索URL search_url = f"https://www.linkedin.com/jobs/search/?keywords={encoded_keyword}&location={encoded_location}" # 调用你的爬虫函数获取数据 job_data = scrape_linkedin_jobs(search_url) # 你的爬虫函数 # 渲染结果页面 return render_template("results.html", jobs=job_data.to_dict("records")) # GET请求时渲染输入表单 return render_template("index.html")
2. 前端表单模板(index.html)
<!DOCTYPE html> <html> <head> <title>LinkedIn职位搜索</title> </head> <body> <h1>LinkedIn职位爬虫</h1> <form method="POST"> <div> <label>职位关键词:</label> <input type="text" name="keyword" placeholder="如Python开发" required> </div> <div> <label>工作地点:</label> <input type="text" name="location" placeholder="如北京" required> </div> <button type="submit">开始搜索</button> </form> </body> </html>
内容的提问来源于stack exchange,提问作者AE13
相关产品推荐
相关产品推荐

