You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy技术问题:移除列表换行符并合并为单行字符串

Fixing Your Scrapy Content Merging Issue

Hey there! I see you're just starting out with Scrapy and Python—great job getting your spider up and running, even with CSV export working! Let's tackle that sitecontent formatting problem you're facing.

First, let's break down the two main issues in your current code:

  1. You're extracting sitecontent as a list of text snippets, but you need a single-line string without newlines.
  2. Using zip() with your fields is causing unexpected behavior (since most fields are single-value lists, while sitecontent is a long list), leading to multiple rows with repeated metadata and only one content snippet each.

Step 1: Merge sitecontent into a Clean Single-Line String

Instead of just extracting the list, we'll clean each text snippet, filter out empty strings, and join everything into one line with single spaces (no newlines). Replace your sitecontent extraction line with this:

# Extract all text snippets from <p> tags, clean each one, filter out empty strings
sitecontent_snippets = response.xpath('//body//p//text()').extract()
# Join into a single line, replacing newlines and multiple spaces with a single space
sitecontent = ' '.join([snippet.strip() for snippet in sitecontent_snippets if snippet.strip()])

This will:

  • Strip leading/trailing whitespace from each snippet
  • Skip any empty snippets (like those that were just newlines or spaces)
  • Join all valid snippets into one string with single spaces between them, no newlines.

Step 2: Fix Your Item Yielding

Right now, zip() is pairing your single-value fields (like sitename, siteurl) with each snippet in sitecontent, which creates way more rows than you want. Instead, you want one row per page with all metadata plus the merged sitecontent.

Also, using [0].getall() is redundant—getall() returns a list, and [0] grabs the first element, but you can just use .get() to get the single string value directly (since those meta tags should only exist once per page).

Here's the revised extraction and item-yielding section:

# Content Extraction - get single string values instead of lists
sitename = response.xpath("//meta[@property='og:site_name']/@content").get()
siteurl = response.xpath("//link[@rel='canonical']/@href").get()
dbCategory = response.xpath("//meta[@name='dbCategory']/@content").get()
titles = response.css('title::text').get()
descriptions = response.xpath("//meta[@name='description']/@content").get()
date = response.xpath("//meta[@name='date']/@content").get()
version = response.xpath("//meta[@name='v']/@content").get()
time = response.xpath("//meta[@name='time']/@content").get()

# Clean and merge sitecontent
sitecontent_snippets = response.xpath('//body//p//text()').extract()
sitecontent = ' '.join([snippet.strip() for snippet in sitecontent_snippets if snippet.strip()])

# Create one item per page, not per content snippet
scraped_info = {
    'sitename': sitename,
    'siteurl': siteurl,
    'dbCategory': dbCategory,
    'title': titles,
    'description': descriptions,
    'date': date,
    'version': version,
    'time': time,
    'sitecontent': sitecontent
}
yield scraped_info

Step 3: Full Revised parse Function

Putting it all together, your parse function will look like this:

def parse(self, response):
    page = response.url.split("/")[-2]
    filename = 'nbtest-%s.html' % page
    with open(filename, 'wb') as f:
        f.write(response.body)
    self.log('Saved file %s' % filename)
    
    # Content Extraction
    print(response.url)
    sitename = response.xpath("//meta[@property='og:site_name']/@content").get()
    siteurl = response.xpath("//link[@rel='canonical']/@href").get()
    dbCategory = response.xpath("//meta[@name='dbCategory']/@content").get()
    titles = response.css('title::text').get()
    descriptions = response.xpath("//meta[@name='description']/@content").get()
    date = response.xpath("//meta[@name='date']/@content").get()
    version = response.xpath("//meta[@name='v']/@content").get()
    time = response.xpath("//meta[@name='time']/@content").get()

    # Clean and merge sitecontent into single line
    sitecontent_snippets = response.xpath('//body//p//text()').extract()
    sitecontent = ' '.join([snippet.strip() for snippet in sitecontent_snippets if snippet.strip()])

    # Yield one item per page
    scraped_info = {
        'sitename': sitename,
        'siteurl': siteurl,
        'dbCategory': dbCategory,
        'title': titles,
        'description': descriptions,
        'date': date,
        'version': version,
        'time': time,
        'sitecontent': sitecontent
    }
    yield scraped_info

    all_pages = response.xpath('//a[contains(@href, "html")]/@href').getall()
    for next_page in all_pages :
        next_page = response.urljoin(next_page)
        yield scrapy.Request(next_page, callback=self.parse)

Why This Works

  • Cleaned Content: We're actively filtering out empty snippets and normalizing whitespace, so you won't get [''] or leftover newlines.
  • Single Item Per Page: Instead of generating multiple rows for one page, you'll have one row with all metadata and the full merged content—perfect for your CSV output.
  • Simpler Extraction: Using .get() instead of [0].getall() gives you direct string values, avoiding unnecessary list handling.

Give this a try, and your CSV should have nice, single-line sitecontent values now!

内容的提问来源于stack exchange,提问作者Armin Abele

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 12:58:11