Scrapy爬虫内存占用过高排查求助(并发设置为3)
Hey there, let's dig into this high memory usage issue you're facing with Scrapy + Splash. Even though you've disabled images in your Lua script and checked via Telnet, that Python process eating up 60%+ memory is definitely frustrating. Here are some potential causes and actionable things to check:
1. Unprocessed Responses or Large HTML Bodies Piling Up
Even with images disabled, the HTML returned by Splash might still be massive—especially since your Lua script waits up to 20 seconds for dynamic content to load. If your spider isn't discarding response objects right after parsing, or if you're storing full HTML content in memory (like in a global list, pipeline, or item field), those can add up fast:
- Double-check your
parsemethod: make sure you're not holding ontoresponse.bodyorresponse.textin persistent variables longer than needed. - If you're using an item pipeline, verify it's not caching all items in memory before writing to a database/file. Try flushing items in small batches instead of hoarding everything in RAM.
2. Accidental Memory Leaks in Global State
Scrapy spiders can quietly accumulate memory if you're using global variables that never get cleared. For example:
- If you have something like
self.collected_items = []in your spider class and keep appending to it without ever emptying it, that's a straight-up memory hog. - Scan your spider, middlewares, and pipelines for any variables that grow over the course of the crawl. Test with a small crawl limit (e.g., 10 requests) and see if memory drops after the crawl finishes—if not, there's likely a leak.
3. Scrapy-Splash Connection Pool Bloat
Even if you suspect the issue is with Scrapy, the integration with Splash could be leaving open connections or unused client objects hanging around in memory:
- Check if you've configured
SPLASH_CONNECTION_TIMEOUTandSPLASH_MAX_CONNECTIONSin your settings. Try settingSPLASH_MAX_CONNECTIONSto match yourCONCURRENT_REQUESTS(6 in your case) to prevent excess, idle connections from cluttering up memory.
4. Telnet's Memory Debugging Has Limits
Telnet's mem command only shows a high-level snapshot of Scrapy's internal memory usage—it won't catch all Python-level memory bloat. To get a clearer picture:
- Use Python's built-in
tracemallocmodule to track exact memory allocations. Add this to your spider's__init__method:
Then, after running the spider for a few minutes, print the top memory consumers:import tracemalloc tracemalloc.start()snapshot = tracemalloc.take_snapshot() top_stats = snapshot.statistics('lineno') print("[Top 10 Memory Hogs]") for stat in top_stats[:10]: print(stat) - Tools like
objgraphcan also help you spot unused objects (e.g., hundreds of orphanedSelectororResponseinstances) that aren't being garbage collected.
5. Garbage Collection Isn't Keeping Up
Python's garbage collector might not clean up short-lived objects (like those created during HTML parsing) as quickly as needed. As a test, try forcing periodic garbage collection:
- Add
import gcat the top of your spider, and in yourparsemethod, callgc.collect()after processing each item. Note: this might hit performance a bit, but it'll tell you if uncollected objects are the culprit.
6. Logging Buffer Overload
You have LOG_FILE enabled, and while writing logs to a file shouldn't use much memory, a large unflushed log buffer could contribute to bloat. Try setting LOG_BUFFER_SIZE to a smaller value (e.g., 100) in your settings to force more frequent flushes to disk.
内容的提问来源于stack exchange,提问作者Milano

