遍历DataFrame中超链接爬取数据时报NoneType无find方法错误如何解决?
错误产生原因
- 节点未加载完成:调用
driver.get()后没有等待页面渲染完成就直接获取源码解析,目标experience-section节点还没生成,soup.find()返回None,后续调用find()方法直接报错 - 目标节点不存在:部分账号的主页本身没有填写工作经历,页面无对应
id的section节点;或者网站前端结构更新,你写的节点选择规则失效,某一步查找返回None - 反爬拦截:高频访问触发网站反爬策略,返回的是验证码页、拦截页而非正常的用户主页,自然找不到目标节点
- 节点层级匹配错误:你代码里的节点层级是硬写的,比如
exp_section.find('ul')之后直接找div再找a,部分页面的工作经历节点层级和你预设的不一样,某一步查找失败返回None
可行修复方案
- 加显式等待:调用
driver.get()后使用Selenium的显式等待逻辑,确认目标experience-section节点渲染完成后再获取页面源码,避免拿未渲染的半成品页面 - 每步增加非空判断:所有节点查找操作后都先判断是否为
None,匹配失败时记录当前出错的profileUrl,跳过该条数据避免程序直接崩溃 - 控制访问频率:每次请求之间加随机延迟,降低被反爬拦截的概率
- 容错取值:工作经历可能有多条,不要硬取下标取元素,避免某条经历缺少对应字段时报错
优化后代码示例
from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By import time import random # 遍历前先设置显式等待最大时长,单位秒 wait = WebDriverWait(driver, 10) for idx, row in df.iterrows(): profile_url = row['profileUrl'] try: driver.get(profile_url) # 等待工作经历板块加载完成 wait.until(EC.presence_of_element_located((By.ID, 'experience-section'))) # 加随机延迟,降低反爬风险 time.sleep(random.uniform(1,3)) src = driver.page_source soup = BeautifulSoup(src, 'html.parser') exp_section = soup.find('section', {'id': 'experience-section'}) if not exp_section: print(f"链接{profile_url}无工作经历板块,跳过") continue ul_node = exp_section.find('ul') if not ul_node: print(f"链接{profile_url}工作经历板块无ul节点,跳过") continue # 遍历所有工作经历条目,不要只取第一条 for li in ul_node.find_all('li', recursive=False): div_node = li.find('div') a_node = div_node.find('a') if div_node else None if not a_node: continue # 每个字段单独判断非空再取值 job_title = a_node.find('h3').get_text().strip() if a_node.find('h3') else '' p_list = a_node.find_all('p') company_name = p_list[1].get_text().strip() if len(p_list)>=2 else '' h4_list = a_node.find_all('h4') joining_date = h4_list[0].find_all('span')[1].get_text().strip() if len(h4_list)>=1 and len(h4_list[0].find_all('span'))>=2 else '' exp = h4_list[1].find_all('span')[1].get_text().strip() if len(h4_list)>=2 and len(h4_list[1].find_all('span'))>=2 else '' # 这里可以把取到的字段存到你需要的地方 print(job_title, company_name, joining_date, exp) except Exception as e: print(f"处理链接{profile_url}出错,错误信息:{str(e)}") continue
内容的提问来源于stack exchange,提问作者Khadija Ewonye Yakubu
相关产品推荐
相关产品推荐

