Scrapy技术问题:移除列表换行符并合并为单行字符串
Hey there! I see you're just starting out with Scrapy and Python—great job getting your spider up and running, even with CSV export working! Let's tackle that sitecontent formatting problem you're facing.
First, let's break down the two main issues in your current code:
- You're extracting
sitecontentas a list of text snippets, but you need a single-line string without newlines. - Using
zip()with your fields is causing unexpected behavior (since most fields are single-value lists, whilesitecontentis a long list), leading to multiple rows with repeated metadata and only one content snippet each.
Step 1: Merge sitecontent into a Clean Single-Line String
Instead of just extracting the list, we'll clean each text snippet, filter out empty strings, and join everything into one line with single spaces (no newlines). Replace your sitecontent extraction line with this:
# Extract all text snippets from <p> tags, clean each one, filter out empty strings sitecontent_snippets = response.xpath('//body//p//text()').extract() # Join into a single line, replacing newlines and multiple spaces with a single space sitecontent = ' '.join([snippet.strip() for snippet in sitecontent_snippets if snippet.strip()])
This will:
- Strip leading/trailing whitespace from each snippet
- Skip any empty snippets (like those that were just newlines or spaces)
- Join all valid snippets into one string with single spaces between them, no newlines.
Step 2: Fix Your Item Yielding
Right now, zip() is pairing your single-value fields (like sitename, siteurl) with each snippet in sitecontent, which creates way more rows than you want. Instead, you want one row per page with all metadata plus the merged sitecontent.
Also, using [0].getall() is redundant—getall() returns a list, and [0] grabs the first element, but you can just use .get() to get the single string value directly (since those meta tags should only exist once per page).
Here's the revised extraction and item-yielding section:
# Content Extraction - get single string values instead of lists sitename = response.xpath("//meta[@property='og:site_name']/@content").get() siteurl = response.xpath("//link[@rel='canonical']/@href").get() dbCategory = response.xpath("//meta[@name='dbCategory']/@content").get() titles = response.css('title::text').get() descriptions = response.xpath("//meta[@name='description']/@content").get() date = response.xpath("//meta[@name='date']/@content").get() version = response.xpath("//meta[@name='v']/@content").get() time = response.xpath("//meta[@name='time']/@content").get() # Clean and merge sitecontent sitecontent_snippets = response.xpath('//body//p//text()').extract() sitecontent = ' '.join([snippet.strip() for snippet in sitecontent_snippets if snippet.strip()]) # Create one item per page, not per content snippet scraped_info = { 'sitename': sitename, 'siteurl': siteurl, 'dbCategory': dbCategory, 'title': titles, 'description': descriptions, 'date': date, 'version': version, 'time': time, 'sitecontent': sitecontent } yield scraped_info
Step 3: Full Revised parse Function
Putting it all together, your parse function will look like this:
def parse(self, response): page = response.url.split("/")[-2] filename = 'nbtest-%s.html' % page with open(filename, 'wb') as f: f.write(response.body) self.log('Saved file %s' % filename) # Content Extraction print(response.url) sitename = response.xpath("//meta[@property='og:site_name']/@content").get() siteurl = response.xpath("//link[@rel='canonical']/@href").get() dbCategory = response.xpath("//meta[@name='dbCategory']/@content").get() titles = response.css('title::text').get() descriptions = response.xpath("//meta[@name='description']/@content").get() date = response.xpath("//meta[@name='date']/@content").get() version = response.xpath("//meta[@name='v']/@content").get() time = response.xpath("//meta[@name='time']/@content").get() # Clean and merge sitecontent into single line sitecontent_snippets = response.xpath('//body//p//text()').extract() sitecontent = ' '.join([snippet.strip() for snippet in sitecontent_snippets if snippet.strip()]) # Yield one item per page scraped_info = { 'sitename': sitename, 'siteurl': siteurl, 'dbCategory': dbCategory, 'title': titles, 'description': descriptions, 'date': date, 'version': version, 'time': time, 'sitecontent': sitecontent } yield scraped_info all_pages = response.xpath('//a[contains(@href, "html")]/@href').getall() for next_page in all_pages : next_page = response.urljoin(next_page) yield scrapy.Request(next_page, callback=self.parse)
Why This Works
- Cleaned Content: We're actively filtering out empty snippets and normalizing whitespace, so you won't get
['']or leftover newlines. - Single Item Per Page: Instead of generating multiple rows for one page, you'll have one row with all metadata and the full merged content—perfect for your CSV output.
- Simpler Extraction: Using
.get()instead of[0].getall()gives you direct string values, avoiding unnecessary list handling.
Give this a try, and your CSV should have nice, single-line sitecontent values now!
内容的提问来源于stack exchange,提问作者Armin Abele

