Scrapy中如何去除或替换<br>标签以清理爬取的文本数据?
Hey there! Great question—you absolutely can fix this directly in Scrapy, no need to wait for Pandas post-processing. Let’s walk through how to get both of your desired outcomes, and why your earlier strip() attempts didn’t work.
First, let’s clarify the issue: when you use //div//text(), Scrapy extracts text from every individual text node inside the <div>. The <br> tag creates empty text nodes around it, which is why you’re getting ['This is an example text.', '', 'This is an example too.']—that empty string is the space between the <br> and the surrounding text. Your strip() calls were only removing specific characters from individual strings, but they weren’t filtering out those empty nodes or combining the text properly.
Option 1: Combine Text into a Single Space-Separate String
This will give you "This is an example text. This is an example too" in your CSV.
Using Raw Scrapy Selectors
Instead of just extracting and stripping, filter out empty strings first, then join the remaining text with spaces:
# Extract all text nodes, strip each one, filter out empty strings text_nodes = [text.strip() for text in response.xpath('//div//text()').extract() if text.strip()] # Join the non-empty nodes into a single string clean_text = ' '.join(text_nodes)
Using ItemLoader (Cleaner for Large Projects)
If you’re using ItemLoader, you can set up input/output processors to handle this automatically:
from scrapy.loader import ItemLoader from scrapy.loader.processors import MapCompose, Join class MyCustomLoader(ItemLoader): # Strip whitespace from every text node default_input_processor = MapCompose(str.strip) # Join all non-empty nodes with a space default_output_processor = Join(' ') # In your spider: loader = MyCustomLoader(item=MyItem(), response=response) loader.add_xpath('content', '//div//text()') item = loader.load_item()
Option 2: Replace <br> with Actual Newlines
This will preserve line breaks in your CSV, giving you:
This is an example text. This is an example too.
For this, you can’t just extract text nodes directly—you need to first replace the <br> tags with newlines, then strip out any remaining HTML tags.
Using Regex with Raw Selectors
import re # Get the full HTML content of the div div_html = response.xpath('//div').get() # Replace <br> (and self-closing <br/>) with newlines text_with_breaks = re.sub(r'<br\s*/?>', '\n', div_html) # Remove all remaining HTML tags and strip extra whitespace clean_text = re.sub(r'<.*?>', '', text_with_breaks).strip()
Using a Custom ItemLoader Processor
Wrap the regex logic into a custom processor for cleaner code:
from scrapy.loader import ItemLoader from scrapy.loader.processors import MapCompose import re def process_br_tags(html): # Replace <br> with newlines html = re.sub(r'<br\s*/?>', '\n', html) # Remove all HTML tags text = re.sub(r'<.*?>', '', html) # Strip leading/trailing whitespace return text.strip() class MyCustomLoader(ItemLoader): # Use the custom processor for the content field content_in = MapCompose(process_br_tags) # In your spider: loader = MyCustomLoader(item=MyItem(), response=response) # Pass the full div HTML instead of just text nodes loader.add_xpath('content', '//div') item = loader.load_item()
Why Your Earlier strip() Calls Didn’t Work
The strip(u'\u0027') and strip(u'\u00A0') calls were targeting specific characters (quotes and non-breaking spaces), but your main issue was empty text nodes from the <br> tag. Those empty strings don’t contain those characters, so stripping them did nothing. The fix is to filter out empty nodes entirely (for Option 1) or process the raw HTML to replace <br> before extracting text (for Option 2).
Both methods work entirely within Scrapy, so you won’t need to use Pandas unless you’re doing additional bulk processing later.
内容的提问来源于stack exchange,提问作者Jarod Santoso

