You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬虫可正常爬取但无法匹配文件关键词且无CSV输出

Fixing Your Scrapy Spider's Missing CSV Output

Got it, let's figure out why your spider's running fine (200s all around) but not spitting out that CSV you need. Here's the breakdown and fix:

The Root Problems

First, let's look at what's going wrong in your current code:

  • Keywords have hidden newlines: When you read lines from dictionarA1.csv, every word includes a trailing \n (newline character). So your XPath is actually searching for things like "a la\n" instead of "a la"—no wonder it's not matching anything on the page!
  • XPath is targeting partial text nodes: Your current XPath //p/text()[contains(..,"{0}")] checks if the parent <p> has the keyword, but only pulls individual text snippets inside the <p>. If the paragraph has any formatting tags (like <b> or <i>), you'll miss parts of the text, or even get nothing at all.

Fixed Spider Code

import scrapy

class StringGrab(scrapy.Spider):
    name = "stringpage"
    start_urls = [
        'https://whatscookingamerica.net/Glossary/A.htm',
    ]

    def parse(self, response):
        # Safely read and clean your keyword list
        with open("dictionarA1.csv", "r", encoding="utf-8") as in_file:
            # Strip newlines/spaces, skip empty lines
            keywords = [line.strip() for line in in_file if line.strip()]
        
        for keyword in keywords:
            # Find all <p> elements that contain the keyword (handle messy whitespace)
            for para in response.xpath(f'//p[contains(normalize-space(.), "{keyword}")]'):
                # Grab the full cleaned text of the paragraph
                full_para_text = para.xpath('normalize-space(.)').get()
                yield {
                    'dictA': full_para_text,
                }

    custom_settings = {
        "DOWNLOAD_DELAY": 1,
        "CONCURRENT_REQUESTS_PER_DOMAIN": 10
    }

What I Changed & Why

  1. Cleaned Up Keyword Input:

    • Used a with statement to open the file (this ensures it gets closed properly, no leftover file handles)
    • Added strip() to yank out newlines, leading/trailing spaces from each keyword—so your XPath searches for the actual term you want
    • Filtered out empty lines to avoid wasting time on blank queries
  2. Fixed the XPath Logic:

    • Now we target the entire <p> element that contains the keyword (using normalize-space(.) to handle extra spaces or line breaks in the paragraph text)
    • Extract the full cleaned text of the paragraph instead of tiny text snippets—so you get the complete relevant content
  3. Removed Deprecated Junk:

    • That sys.setdefaultencoding('utf8') code is deprecated in Python 3 and will throw errors. Scrapy handles UTF-8 by default, so we don't need it at all.

How to Get Your CSV

To generate the CSV output, just run your spider with this command:

scrapy crawl stringpage -o output.csv

This will create a properly formatted CSV file with all the matching paragraphs you're after.

内容的提问来源于stack exchange,提问作者Schneejäger

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 08:12:52