Scrapy爬虫可正常爬取但无法匹配文件关键词且无CSV输出
Fixing Your Scrapy Spider's Missing CSV Output
Got it, let's figure out why your spider's running fine (200s all around) but not spitting out that CSV you need. Here's the breakdown and fix:
The Root Problems
First, let's look at what's going wrong in your current code:
- Keywords have hidden newlines: When you read lines from
dictionarA1.csv, everywordincludes a trailing\n(newline character). So your XPath is actually searching for things like"a la\n"instead of"a la"—no wonder it's not matching anything on the page! - XPath is targeting partial text nodes: Your current XPath
//p/text()[contains(..,"{0}")]checks if the parent<p>has the keyword, but only pulls individual text snippets inside the<p>. If the paragraph has any formatting tags (like<b>or<i>), you'll miss parts of the text, or even get nothing at all.
Fixed Spider Code
import scrapy class StringGrab(scrapy.Spider): name = "stringpage" start_urls = [ 'https://whatscookingamerica.net/Glossary/A.htm', ] def parse(self, response): # Safely read and clean your keyword list with open("dictionarA1.csv", "r", encoding="utf-8") as in_file: # Strip newlines/spaces, skip empty lines keywords = [line.strip() for line in in_file if line.strip()] for keyword in keywords: # Find all <p> elements that contain the keyword (handle messy whitespace) for para in response.xpath(f'//p[contains(normalize-space(.), "{keyword}")]'): # Grab the full cleaned text of the paragraph full_para_text = para.xpath('normalize-space(.)').get() yield { 'dictA': full_para_text, } custom_settings = { "DOWNLOAD_DELAY": 1, "CONCURRENT_REQUESTS_PER_DOMAIN": 10 }
What I Changed & Why
Cleaned Up Keyword Input:
- Used a
withstatement to open the file (this ensures it gets closed properly, no leftover file handles) - Added
strip()to yank out newlines, leading/trailing spaces from each keyword—so your XPath searches for the actual term you want - Filtered out empty lines to avoid wasting time on blank queries
- Used a
Fixed the XPath Logic:
- Now we target the entire
<p>element that contains the keyword (usingnormalize-space(.)to handle extra spaces or line breaks in the paragraph text) - Extract the full cleaned text of the paragraph instead of tiny text snippets—so you get the complete relevant content
- Now we target the entire
Removed Deprecated Junk:
- That
sys.setdefaultencoding('utf8')code is deprecated in Python 3 and will throw errors. Scrapy handles UTF-8 by default, so we don't need it at all.
- That
How to Get Your CSV
To generate the CSV output, just run your spider with this command:
scrapy crawl stringpage -o output.csv
This will create a properly formatted CSV file with all the matching paragraphs you're after.
内容的提问来源于stack exchange,提问作者Schneejäger
相关产品推荐
相关产品推荐

