You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy中如何去除或替换<br>标签以清理爬取的文本数据?

Handling <br> Tags in Scrapy for Clean CSV Output

Hey there! Great question—you absolutely can fix this directly in Scrapy, no need to wait for Pandas post-processing. Let’s walk through how to get both of your desired outcomes, and why your earlier strip() attempts didn’t work.

First, let’s clarify the issue: when you use //div//text(), Scrapy extracts text from every individual text node inside the <div>. The <br> tag creates empty text nodes around it, which is why you’re getting ['This is an example text.', '', 'This is an example too.']—that empty string is the space between the <br> and the surrounding text. Your strip() calls were only removing specific characters from individual strings, but they weren’t filtering out those empty nodes or combining the text properly.

Option 1: Combine Text into a Single Space-Separate String

This will give you "This is an example text. This is an example too" in your CSV.

Using Raw Scrapy Selectors

Instead of just extracting and stripping, filter out empty strings first, then join the remaining text with spaces:

# Extract all text nodes, strip each one, filter out empty strings
text_nodes = [text.strip() for text in response.xpath('//div//text()').extract() if text.strip()]
# Join the non-empty nodes into a single string
clean_text = ' '.join(text_nodes)

Using ItemLoader (Cleaner for Large Projects)

If you’re using ItemLoader, you can set up input/output processors to handle this automatically:

from scrapy.loader import ItemLoader
from scrapy.loader.processors import MapCompose, Join

class MyCustomLoader(ItemLoader):
    # Strip whitespace from every text node
    default_input_processor = MapCompose(str.strip)
    # Join all non-empty nodes with a space
    default_output_processor = Join(' ')

# In your spider:
loader = MyCustomLoader(item=MyItem(), response=response)
loader.add_xpath('content', '//div//text()')
item = loader.load_item()

Option 2: Replace <br> with Actual Newlines

This will preserve line breaks in your CSV, giving you:

This is an example text.
This is an example too.

For this, you can’t just extract text nodes directly—you need to first replace the <br> tags with newlines, then strip out any remaining HTML tags.

Using Regex with Raw Selectors

import re

# Get the full HTML content of the div
div_html = response.xpath('//div').get()
# Replace <br> (and self-closing <br/>) with newlines
text_with_breaks = re.sub(r'<br\s*/?>', '\n', div_html)
# Remove all remaining HTML tags and strip extra whitespace
clean_text = re.sub(r'<.*?>', '', text_with_breaks).strip()

Using a Custom ItemLoader Processor

Wrap the regex logic into a custom processor for cleaner code:

from scrapy.loader import ItemLoader
from scrapy.loader.processors import MapCompose
import re

def process_br_tags(html):
    # Replace <br> with newlines
    html = re.sub(r'<br\s*/?>', '\n', html)
    # Remove all HTML tags
    text = re.sub(r'<.*?>', '', html)
    # Strip leading/trailing whitespace
    return text.strip()

class MyCustomLoader(ItemLoader):
    # Use the custom processor for the content field
    content_in = MapCompose(process_br_tags)

# In your spider:
loader = MyCustomLoader(item=MyItem(), response=response)
# Pass the full div HTML instead of just text nodes
loader.add_xpath('content', '//div')
item = loader.load_item()

Why Your Earlier strip() Calls Didn’t Work

The strip(u'\u0027') and strip(u'\u00A0') calls were targeting specific characters (quotes and non-breaking spaces), but your main issue was empty text nodes from the <br> tag. Those empty strings don’t contain those characters, so stripping them did nothing. The fix is to filter out empty nodes entirely (for Option 1) or process the raw HTML to replace <br> before extracting text (for Option 2).

Both methods work entirely within Scrapy, so you won’t need to use Pandas unless you’re doing additional bulk processing later.

内容的提问来源于stack exchange,提问作者Jarod Santoso

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 04:12:39