Python结合BeautifulSoup/Selenium遍历ul下li提取数据报错问题
问题根因
报错来自三个写法错误:
find_all()返回的是ResultSet类型,本质是匹配到的所有标签组成的列表,你在循环里对整个列表对象调用find()方法,而不是对当前遍历到的单个li元素调用,自然会抛出属性不存在的错误。- 直接写
for li in searchList遍历ul标签是错的:单个BeautifulSoup标签对象直接遍历会返回所有子节点,包括标签之间的换行、空白文本、注释节点,不会自动筛选出li元素。 - 你代码里部分class属性值因为手动换行多了多余空格、甚至断成两截,会导致BeautifulSoup匹配不到对应标签,比如ul的class值、location对应的class值都有换行导致的格式错误。
修正实现逻辑
- 定位到目标ul标签后,调用
find_all()筛选出ul下所有符合class规则的li元素,存为li集合,不要提前用find()只取第一个li。 - 遍历li集合时,用循环拿到的单个li对象查找内部的姓名、职位等字段,绝对不能用存整个li集合的列表变量调用
find()。 - 每个字段查找后加空值判断,避免列表里混入广告、占位项导致找不到标签时抛出
NoneType错误。
修正后可运行代码
from bs4 import BeautifulSoup page_source = driver.page_source soup = BeautifulSoup(page_source, features='html.parser') searchResCon = soup.find('div', {'class':'search-results-container'}) followerCol = searchResCon.find('div', {'class':'ph0 pv2 artdeco-card mb2'}) # 定位目标ul,注意class值不要断行、不要加多余空格 searchList = followerCol.find('ul', {'class':'reusable-search__entity-result-list list-style-none'}) # 取出ul下所有符合条件的li all_li_items = searchList.find_all('li', {'class':'reusable-search__result-container'}) # 可以先打印长度确认是否拿到至少10个li # print(len(all_li_items)) contactsInfo = [] for singleCon in all_li_items: # 每个字段做空值兼容,避免缺字段报错 name_tag = singleCon.find('span', {'aria-hidden':'true'}) name = name_tag.get_text().strip() if name_tag else "无姓名信息" title_tag = singleCon.find('div', {'class':'entity-result__primary-subtitle t-14 t-black t-normal'}) title = title_tag.get_text().strip() if title_tag else "无职位信息" location_tag = singleCon.find('div', {'class':'entity-result__secondary-subtitle t-14 t-normal'}) location = location_tag.get_text().strip() if location_tag else "无所在地信息" hashtag_tag = singleCon.find('p', {'class':'entity-result__summary entity-result__summary--2-lines t-12 t-black--light mb1'}) hashtag = hashtag_tag.get_text().strip() if hashtag_tag else "无简介标签信息" follower_tag = singleCon.find('span',{'class':'entity-result__simple-insight-text entity-result__simple-insight-text--small'}) follower = follower_tag.get_text().strip() if follower_tag else "无关注/共同好友信息" # 写入结果列表 contactsInfo.append('-' * 30 + '\n') contactsInfo.append('-' * 30 + '\n') contactsInfo.append(f'Name: {name}\n') contactsInfo.append(f'Title: {title}\n') contactsInfo.append(f'Location: {location}\n') contactsInfo.append(f'Hashtag: {hashtag}\n') contactsInfo.append(f'Follower & Mutual: {follower}\n')
额外注意点
- 所有标签的class、id等属性值写的时候不要手动换行,避免引入多余空格导致匹配失败。
- 如果
len(all_li_items)返回的数量和你预期的10个不符,往上检查每一步的标签定位是否正确,可以逐段打印标签的prettify()结果确认定位到的内容是否正确。 - 要是页面是动态加载的,确认你取
page_source的时候所有li已经加载完成,不然会出现拿到的li数量不足的问题。
内容的提问来源于stack exchange,提问作者alneckoc
相关产品推荐
相关产品推荐

