Scrapy中如何获取div元素下的完整文本内容?
Hey there! I've run into this exact issue before when scraping content with nested elements, so I know how frustrating it can be to miss out on text inside child tags like <p>, <a>, or <ul>. Let me walk you through the best ways to grab the complete, formatted text from your usefulInfo divs.
Solution 1: Use //text() to Grab All Descendant Text Nodes
This method lets you fetch every text node under the parent element, then combine them into a single string. It gives you full control over how you handle whitespace.
In your Spider, you'd do something like this:
def parse(self, response): # Iterate over each usefulInfo div for useful_div in response.xpath('//div[@class="usefulInfo"]'): # Get all text nodes inside the div (including nested elements) text_list = useful_div.xpath('.//text()').getall() # Join the list into a single string, stripping extra whitespace full_text = ' '.join(text.strip() for text in text_list if text.strip()) # Now full_text contains all the text from the div and its children yield {'full_content': full_text}
- The
.//text()XPath selects all text nodes under the currentuseful_div(the dot ensures we only look inside this specific div, not the whole page). - We use a generator expression to strip each text segment and skip empty strings, then join everything with spaces for clean formatting.
Solution 2: Use normalize-space() for Clean, Concise Text
If you want to automatically handle messy whitespace (like multiple newlines or spaces), XPath's normalize-space() function is perfect. It collapses all consecutive whitespace into single spaces and trims the ends.
Here's how to use it:
def parse(self, response): for useful_div in response.xpath('//div[@class="usefulInfo"]'): # Extract and normalize all text in the div full_text = useful_div.xpath('normalize-space(.)').get() yield {'full_content': full_text}
- The
.innormalize-space(.)refers to the currentuseful_divelement. - This returns a single, clean string with all nested text included—no need to manually join or filter.
Bonus: Preserve Paragraph/List Structure
If you want the text to match the page's visual structure (e.g., newlines for paragraphs or lists), you can iterate through the direct child elements and add formatting:
def parse(self, response): for useful_div in response.xpath('//div[@class="usefulInfo"]'): full_text = "" # Loop through direct child elements of the div for child in useful_div.xpath('./*'): tag_name = child.xpath('name()').get() # Get text from this child and its nested elements child_text = ' '.join(child.xpath('.//text()').getall()).strip() if child_text: # Add a newline before paragraphs or lists to mimic page structure if tag_name in ['p', 'ul']: full_text += f"\n{child_text}" else: full_text += f" {child_text}" # Trim any leading/trailing spaces full_text = full_text.strip() yield {'full_content': full_text}
This will give you text that looks closer to how it appears on the page, with breaks between paragraphs and lists.
For your specific HTML structure, both Solution 1 and 2 will capture text like "Some text Useful paragraph Useful paragraph with Important Keywords Some interesting data" (with clean whitespace), which matches exactly what's displayed on the page.
内容的提问来源于stack exchange,提问作者Nicolas Gaitan

