Scrapy选择器extract()与extract_first()方法返回值不一致问题咨询
Hey there! Let's break down why extract() and extract_first() give you different results in your Scrapy code, and how to fix it based on what you need.
Core Differences Between the Two Methods
First, let's clarify what each method does—this is the root of your differing results:
extract(): Pulls text from all matching selector nodes and returns a list of strings. Even if only one result is found, it'll still wrap it in a list. If no matches exist, it returns an empty list[].extract_first(): Only pulls text from the first matching selector node and returns a single string. If no matches are found, it defaults toNone(you can set a custom fallback withextract_first(default="your_default_value")).
Why Your Code Returns Different Results
Looking at your code snippets, let's use a real-world example to illustrate:
Suppose your target HTML looks like this:
<div class="col-sm-7 col-md-9"> <p class="movie__option"> 复仇者联盟4 <span>时长: 181分钟</span> </p> <p class="movie__option"> 国语版 <span>IMAX厅</span> </p> </div>
- If you use
i.xpath("text()").extract()(looping through each<p>node), you'll get lists like['复仇者联盟4', '\n ']and['国语版', '\n ']—it captures all text nodes, including whitespace. - If you switch to
i.xpath("text()").extract_first(), you'll get single strings:'复仇者联盟4'and'国语版'—it only grabs the first text node from each<p>.
If you run these methods outside a loop (e.g., directly on the div selector), the gap gets even bigger:
div.xpath("text()").extract()returns a combined list of all text nodes from every matching<p>.div.xpath("text()").extract_first()returns only the very first text node from the entire set of matches.
Fixes & Best Practices
Choose the method that aligns with your end goal:
If you need all matching text (as a cleaned list)
Stick withextract()(or its more intuitive aliasgetall()), but filter out empty/whitespace-only strings:def parse_with_cleaned_extract(self, response): p_nodes = response.xpath("//div[@class='col-sm-7 col-md-9']/p[@class='movie__option']") all_data = [] for node in p_nodes: # Extract and clean text cleaned_texts = [t.strip() for t in node.xpath("text()").getall() if t.strip()] all_data.extend(cleaned_texts) return all_dataIf you only need the first matching text
Useextract_first()(or its aliasget()) and set a fallback to avoidNonevalues:def parse_with_extract_first(self, response): first_valid_text = response.xpath("//div[@class='col-sm-7 col-md-9']/p[@class='movie__option']/text()").get(default="No data found") return first_valid_textPro Tip: Use
get()andgetall()
Scrapy added these aliases for readability—get()=extract_first(),getall()=extract(). They make your code easier to understand at a glance.
内容的提问来源于stack exchange,提问作者6ones

