You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

遍历DataFrame中超链接爬取数据时报NoneType无find方法错误如何解决?

错误产生原因
  • 节点未加载完成:调用driver.get()后没有等待页面渲染完成就直接获取源码解析,目标experience-section节点还没生成,soup.find()返回None,后续调用find()方法直接报错
  • 目标节点不存在:部分账号的主页本身没有填写工作经历,页面无对应id的section节点;或者网站前端结构更新,你写的节点选择规则失效,某一步查找返回None
  • 反爬拦截:高频访问触发网站反爬策略,返回的是验证码页、拦截页而非正常的用户主页,自然找不到目标节点
  • 节点层级匹配错误:你代码里的节点层级是硬写的,比如exp_section.find('ul')之后直接找div再找a,部分页面的工作经历节点层级和你预设的不一样,某一步查找失败返回None
可行修复方案
  • 加显式等待:调用driver.get()后使用Selenium的显式等待逻辑,确认目标experience-section节点渲染完成后再获取页面源码,避免拿未渲染的半成品页面
  • 每步增加非空判断:所有节点查找操作后都先判断是否为None,匹配失败时记录当前出错的profileUrl,跳过该条数据避免程序直接崩溃
  • 控制访问频率:每次请求之间加随机延迟,降低被反爬拦截的概率
  • 容错取值:工作经历可能有多条,不要硬取下标取元素,避免某条经历缺少对应字段时报错
优化后代码示例
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By
import time
import random

# 遍历前先设置显式等待最大时长,单位秒
wait = WebDriverWait(driver, 10)

for idx, row in df.iterrows():
    profile_url = row['profileUrl']
    try:
        driver.get(profile_url)
        # 等待工作经历板块加载完成
        wait.until(EC.presence_of_element_located((By.ID, 'experience-section')))
        # 加随机延迟,降低反爬风险
        time.sleep(random.uniform(1,3))
        
        src = driver.page_source
        soup = BeautifulSoup(src, 'html.parser')
        exp_section = soup.find('section', {'id': 'experience-section'})
        if not exp_section:
            print(f"链接{profile_url}无工作经历板块,跳过")
            continue
        
        ul_node = exp_section.find('ul')
        if not ul_node:
            print(f"链接{profile_url}工作经历板块无ul节点,跳过")
            continue
        
        # 遍历所有工作经历条目,不要只取第一条
        for li in ul_node.find_all('li', recursive=False):
            div_node = li.find('div')
            a_node = div_node.find('a') if div_node else None
            if not a_node:
                continue
            
            # 每个字段单独判断非空再取值
            job_title = a_node.find('h3').get_text().strip() if a_node.find('h3') else ''
            p_list = a_node.find_all('p')
            company_name = p_list[1].get_text().strip() if len(p_list)>=2 else ''
            h4_list = a_node.find_all('h4')
            joining_date = h4_list[0].find_all('span')[1].get_text().strip() if len(h4_list)>=1 and len(h4_list[0].find_all('span'))>=2 else ''
            exp = h4_list[1].find_all('span')[1].get_text().strip() if len(h4_list)>=2 and len(h4_list[1].find_all('span'))>=2 else ''
            
            # 这里可以把取到的字段存到你需要的地方
            print(job_title, company_name, joining_date, exp)

    except Exception as e:
        print(f"处理链接{profile_url}出错,错误信息:{str(e)}")
        continue

内容的提问来源于stack exchange,提问作者Khadija Ewonye Yakubu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 00:15:00