You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup find_all()爬取时返回多余元素的问题求助

解决BeautifulSoup匹配单个class时包含多class元素的问题

当你用class_='list-row'查找标签时,BeautifulSoup会匹配所有包含该class的元素,不管它有没有其他额外class。要只匹配仅含list-row这一个class的<li>标签,可以用以下几种方法:

方法1:使用CSS精确属性选择器

用select()方法配合[class="list-row"],严格匹配class属性值完全等于list-row的元素:

with open('index.html', 'r') as f:
    contents = f.read()
soup = BeautifulSoup(contents, "html.parser")  # 修正:原代码误将contents写成html
main_block = soup.find('ul', class_='list')  # 修正:原代码的conn应为soup的笔误
for li in main_block.select('li[class="list-row"]'):
    print(li.prettify())

方法2:用lambda表达式过滤

在find_all()里传入自定义判断逻辑,检查class列表是否只有list-row:

with open('index.html', 'r') as f:
    contents = f.read()
soup = BeautifulSoup(contents, "html.parser")
main_block = soup.find('ul', class_='list')
# 过滤class属性恰好是['list-row']的li标签
for li in main_block.find_all('li', class_=lambda x: x == 'list-row'):
    print(li.prettify())

原代码的小问题修正

  • 你定义了contents = f.read()但未使用,soup初始化时错误引用了不存在的html变量
  • 原代码里的conn(limit_txt,limit)应为soup的笔误,否则会触发未定义变量错误

内容的提问来源于stack exchange,提问作者treboris

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 06:30:11