You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python结合BeautifulSoup/Selenium遍历ul下li提取数据报错问题

问题根因

报错来自三个写法错误:

  • find_all()返回的是ResultSet类型,本质是匹配到的所有标签组成的列表,你在循环里对整个列表对象调用find()方法,而不是对当前遍历到的单个li元素调用,自然会抛出属性不存在的错误。
  • 直接写for li in searchList遍历ul标签是错的:单个BeautifulSoup标签对象直接遍历会返回所有子节点,包括标签之间的换行、空白文本、注释节点,不会自动筛选出li元素。
  • 你代码里部分class属性值因为手动换行多了多余空格、甚至断成两截,会导致BeautifulSoup匹配不到对应标签,比如ul的class值、location对应的class值都有换行导致的格式错误。
修正实现逻辑
  1. 定位到目标ul标签后,调用find_all()筛选出ul下所有符合class规则的li元素,存为li集合,不要提前用find()只取第一个li。
  2. 遍历li集合时,用循环拿到的单个li对象查找内部的姓名、职位等字段,绝对不能用存整个li集合的列表变量调用find()。
  3. 每个字段查找后加空值判断,避免列表里混入广告、占位项导致找不到标签时抛出NoneType错误。
修正后可运行代码
from bs4 import BeautifulSoup

page_source = driver.page_source
soup = BeautifulSoup(page_source, features='html.parser')
searchResCon = soup.find('div', {'class':'search-results-container'})
followerCol = searchResCon.find('div', {'class':'ph0 pv2 artdeco-card mb2'})
# 定位目标ul,注意class值不要断行、不要加多余空格
searchList = followerCol.find('ul', {'class':'reusable-search__entity-result-list list-style-none'})
# 取出ul下所有符合条件的li
all_li_items = searchList.find_all('li', {'class':'reusable-search__result-container'})

# 可以先打印长度确认是否拿到至少10个li
# print(len(all_li_items))

contactsInfo = []
for singleCon in all_li_items:
    # 每个字段做空值兼容,避免缺字段报错
    name_tag = singleCon.find('span', {'aria-hidden':'true'})
    name = name_tag.get_text().strip() if name_tag else "无姓名信息"

    title_tag = singleCon.find('div', {'class':'entity-result__primary-subtitle t-14 t-black t-normal'})
    title = title_tag.get_text().strip() if title_tag else "无职位信息"

    location_tag = singleCon.find('div', {'class':'entity-result__secondary-subtitle t-14 t-normal'})
    location = location_tag.get_text().strip() if location_tag else "无所在地信息"

    hashtag_tag = singleCon.find('p', {'class':'entity-result__summary entity-result__summary--2-lines t-12 t-black--light mb1'})
    hashtag = hashtag_tag.get_text().strip() if hashtag_tag else "无简介标签信息"

    follower_tag = singleCon.find('span',{'class':'entity-result__simple-insight-text entity-result__simple-insight-text--small'})
    follower = follower_tag.get_text().strip() if follower_tag else "无关注/共同好友信息"

    # 写入结果列表
    contactsInfo.append('-' * 30 + '\n')
    contactsInfo.append('-' * 30 + '\n')
    contactsInfo.append(f'Name: {name}\n')
    contactsInfo.append(f'Title: {title}\n')
    contactsInfo.append(f'Location: {location}\n')
    contactsInfo.append(f'Hashtag: {hashtag}\n')
    contactsInfo.append(f'Follower & Mutual: {follower}\n')
额外注意点
  • 所有标签的class、id等属性值写的时候不要手动换行,避免引入多余空格导致匹配失败。
  • 如果len(all_li_items)返回的数量和你预期的10个不符,往上检查每一步的标签定位是否正确,可以逐段打印标签的prettify()结果确认定位到的内容是否正确。
  • 要是页面是动态加载的,确认你取page_source的时候所有li已经加载完成,不然会出现拿到的li数量不足的问题。

内容的提问来源于stack exchange,提问作者alneckoc

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.30 08:27:20