You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy中如何获取div元素下的完整文本内容?

How to Extract Full Text from Parent Elements in Scrapy

Hey there! I've run into this exact issue before when scraping content with nested elements, so I know how frustrating it can be to miss out on text inside child tags like <p>, <a>, or <ul>. Let me walk you through the best ways to grab the complete, formatted text from your usefulInfo divs.

Solution 1: Use //text() to Grab All Descendant Text Nodes

This method lets you fetch every text node under the parent element, then combine them into a single string. It gives you full control over how you handle whitespace.

In your Spider, you'd do something like this:

def parse(self, response):
    # Iterate over each usefulInfo div
    for useful_div in response.xpath('//div[@class="usefulInfo"]'):
        # Get all text nodes inside the div (including nested elements)
        text_list = useful_div.xpath('.//text()').getall()
        # Join the list into a single string, stripping extra whitespace
        full_text = ' '.join(text.strip() for text in text_list if text.strip())
        # Now full_text contains all the text from the div and its children
        yield {'full_content': full_text}
  • The .//text() XPath selects all text nodes under the current useful_div (the dot ensures we only look inside this specific div, not the whole page).
  • We use a generator expression to strip each text segment and skip empty strings, then join everything with spaces for clean formatting.

Solution 2: Use normalize-space() for Clean, Concise Text

If you want to automatically handle messy whitespace (like multiple newlines or spaces), XPath's normalize-space() function is perfect. It collapses all consecutive whitespace into single spaces and trims the ends.

Here's how to use it:

def parse(self, response):
    for useful_div in response.xpath('//div[@class="usefulInfo"]'):
        # Extract and normalize all text in the div
        full_text = useful_div.xpath('normalize-space(.)').get()
        yield {'full_content': full_text}
  • The . in normalize-space(.) refers to the current useful_div element.
  • This returns a single, clean string with all nested text included—no need to manually join or filter.

Bonus: Preserve Paragraph/List Structure

If you want the text to match the page's visual structure (e.g., newlines for paragraphs or lists), you can iterate through the direct child elements and add formatting:

def parse(self, response):
    for useful_div in response.xpath('//div[@class="usefulInfo"]'):
        full_text = ""
        # Loop through direct child elements of the div
        for child in useful_div.xpath('./*'):
            tag_name = child.xpath('name()').get()
            # Get text from this child and its nested elements
            child_text = ' '.join(child.xpath('.//text()').getall()).strip()
            if child_text:
                # Add a newline before paragraphs or lists to mimic page structure
                if tag_name in ['p', 'ul']:
                    full_text += f"\n{child_text}"
                else:
                    full_text += f" {child_text}"
        # Trim any leading/trailing spaces
        full_text = full_text.strip()
        yield {'full_content': full_text}

This will give you text that looks closer to how it appears on the page, with breaks between paragraphs and lists.

For your specific HTML structure, both Solution 1 and 2 will capture text like "Some text Useful paragraph Useful paragraph with Important Keywords Some interesting data" (with clean whitespace), which matches exactly what's displayed on the page.

内容的提问来源于stack exchange,提问作者Nicolas Gaitan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 10:22:49