You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Beautiful Soup爬取Indeed网站返回结果为0的问题排查

问题排查与解决:Indeed爬虫返回0的原因

核心问题1:请求头参数错误

你的headers字典里的键写成了User_Agent,正确的浏览器请求头键名是User-Agent(下划线需改为横杠)。这个错误会让服务器识别请求为非浏览器发起,返回的页面结构和你预期的完全不符,自然找不到目标class的元素。

修正后的extract函数请求头部分:

headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:130.0) Gecko/20100101 Firefox/130.0'}

核心问题2:页面class可能已更新

Indeed的页面结构会频繁变动,你截图里的cardOutline tapItem dd-privacy-allow组合class可能已经失效。建议在extract函数里添加一行代码,打印返回页面的前1000个字符,确认实际的HTML结构:

print(soup.prettify()[:1000])

根据实际返回的内容,重新定位职位卡片的class(比如当前Indeed常用job_seen_beacon作为职位卡片的主class)。

额外注意:反爬机制限制

即使修正请求头,也可能遇到Indeed的反爬策略,返回验证码页面或空内容。这种情况下可以尝试:

  • 从浏览器复制真实的Cookie参数添加到请求头
  • 使用requests.Session()维持会话
  • 若页面有动态加载内容,改用Selenium等工具模拟浏览器行为

修改后的完整示例代码

import requests
from bs4 import BeautifulSoup

def extract(page):
    headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:130.0) Gecko/20100101 Firefox/130.0'}
    url = f'https://www.indeed.com/jobs?q=data+analyst&start={page}'
    r = requests.get(url, headers=headers)
    print(f"请求状态码: {r.status_code}") # 验证请求是否成功
    soup = BeautifulSoup(r.content, 'html.parser')
    # print(soup.prettify()[:1000]) # 查看实际页面结构
    return soup

def transform(soup):
    # 根据实际页面结构调整class选择器
    divs = soup.find_all('div', class_='job_seen_beacon')
    return len(divs)

c = extract(0)
print(transform(c))

内容的提问来源于stack exchange,提问作者Jonathan Mporampora

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 18:22:33