Scrapy如何选取无属性的div元素并提取文本?
Got it, let's figure out how to grab those attribute-free divs you need! The issue with a plain div selector is that it picks up every div—including the ones with classes like inside or hello. We need a way to filter out any div that has any attribute at all.
The Best Solution: Use XPath Selector
Scrapy's XPath support makes this super straightforward. You can use the not(@*) condition to target divs with zero attributes:
# In your Scrapy spider, use this selector target_texts = response.xpath('//div[not(@*)]/text()').getall()
Let's break this down:
//div: Finds all div elements anywhere in the HTML[not(@*)]: Filters out any div that has any attribute (class, id, style, etc.)/text(): Extracts the text content from the matching divsgetall(): Returns all matching text values as a list, which will be exactly['test', 'test3', 'test5', 'test6']for your sample HTML.
Alternative: CSS Selector (Less Ideal for "No Attributes" Case)
CSS doesn't have a direct way to check for "no attributes" like XPath does. You could chain :not() pseudo-classes to exclude divs with specific attributes, but this isn't foolproof (it won't catch rare/unexpected attributes):
# Only excludes divs with class, id, or style—won't catch all possible attributes target_texts = response.css('div:not([class]):not([id]):not([style])::text').getall()
For your exact sample HTML, this would work, but XPath is the better choice if you need to strictly target divs with zero attributes of any kind.
内容的提问来源于stack exchange,提问作者Mernayi

