You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python使用BeautifulSoup遍历网页分页无法获取下一页按钮href如何解决

修复方案

你的代码存在3个核心问题,按优先级调整后即可正常获取分页地址:

1. 调整下一页逻辑的缩进位置

你当前把获取下一页地址的代码写在了遍历职位卡片的for循环内部,这会导致每处理1条职位数据就重复查询一次下一页按钮,还可能出现只处理了部分职位就提前终止分页的问题。需要把try-except块移动到和for card in cards同级的位置,等当前页所有职位都解析完成后再切换下一页。

2. 给请求添加UA头规避反爬

Indeed对无标识的爬虫请求拦截率极高,默认的requests请求头会被直接识别拦截,返回的异常页面里不存在下一页按钮,自然拿不到href属性。发起请求时添加正常浏览器的User-Agent即可解决大部分拦截问题。

3. 优化下一页选择器的兼容性

如果站点更新了下一页按钮的aria-label属性值,可以新增按类名查询的兜底逻辑,提高代码健壮性。

修复后完整代码

import csv
from datetime import datetime
import requests
from bs4 import BeautifulSoup

# 新增请求头,替换为你自己浏览器的UA也可以
HEADERS = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
}

def get_url(position, location):
    "Generate a url from position and location"
    template = 'https://mx.indeed.com/jobs?q={}&l={}'
    url = template.format(position, location)
    return url

def get_record(card):
    spantag = card.h2.span
    job_title = spantag.get('title')
    job_url = 'https://www.indeed.com' + card.get('href')
    company = card.find('span', 'companyName').text
    job_location = card.find('div', 'companyLocation').text
    job_summary = card.find('div', 'job-snippet').text.strip()
    post_date = card.find('span', 'date').text
    today = datetime.today().strftime('%Y-%m-%d')
    
    record = (job_title, company, job_location, post_date, today, job_summary, job_url)
    
    return record

def main(position, location):
    records = []
    url = get_url(position, location)
    
    while True:
        # 请求时添加headers
        response = requests.get(url, headers=HEADERS)
        soup = BeautifulSoup(response.text, 'html.parser')
        cards = soup.find_all('a', 'tapItem')

        for card in cards:
            record = get_record(card)
            records.append(record)

        # 调整缩进:当前页所有内容处理完成后再找下一页
        try:
            # 可以先按aria-label找,找不到再按类名兜底
            next_btn = soup.find('a', {'aria-label':'Siguiente »'}) or soup.find('a', {'class': 'pagination-next'})
            url = 'https://mx.indeed.com' + next_btn.get('href')
        except AttributeError:
            # 找不到下一页按钮说明到最后一页,终止循环
            break
    
    with open('results_Indeed.csv', 'w', newline = '', encoding = 'utf-8') as f:
        writer = csv.writer(f)
        writer.writerow(['JobTitle', 'Company', 'Location', 'PostDate', 'Date', 'Summary', 'URL'])
        writer.writerows(records)

# 测试运行示例,替换成你要搜索的职位和地点即可
# main("python", "Mexico City")

内容的提问来源于stack exchange,提问作者Benjamin Lopez

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 04:09:04