You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy如何选取无属性的div元素并提取文本?

How to Select Divs with No Attributes in Scrapy

Got it, let's figure out how to grab those attribute-free divs you need! The issue with a plain div selector is that it picks up every div—including the ones with classes like inside or hello. We need a way to filter out any div that has any attribute at all.

The Best Solution: Use XPath Selector

Scrapy's XPath support makes this super straightforward. You can use the not(@*) condition to target divs with zero attributes:

# In your Scrapy spider, use this selector
target_texts = response.xpath('//div[not(@*)]/text()').getall()

Let's break this down:

  • //div: Finds all div elements anywhere in the HTML
  • [not(@*)]: Filters out any div that has any attribute (class, id, style, etc.)
  • /text(): Extracts the text content from the matching divs
  • getall(): Returns all matching text values as a list, which will be exactly ['test', 'test3', 'test5', 'test6'] for your sample HTML.

Alternative: CSS Selector (Less Ideal for "No Attributes" Case)

CSS doesn't have a direct way to check for "no attributes" like XPath does. You could chain :not() pseudo-classes to exclude divs with specific attributes, but this isn't foolproof (it won't catch rare/unexpected attributes):

# Only excludes divs with class, id, or style—won't catch all possible attributes
target_texts = response.css('div:not([class]):not([id]):not([style])::text').getall()

For your exact sample HTML, this would work, but XPath is the better choice if you need to strictly target divs with zero attributes of any kind.

内容的提问来源于stack exchange,提问作者Mernayi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 04:03:10