Scrapy爬取影院排片时触发KeyError,求助排查问题
Let's break down why you're hitting that KeyError and how to fix it step by step:
First, Understand the Root Cause
A KeyError here means your Pipeline is trying to access fields like time, date, or title that don't exist in the ShowItem being passed in. This almost always happens because those fields weren't successfully populated by your spider—either the XPath queries are wrong, or the spider isn't processing items correctly.
1. Fix the Spider's Loop (Critical Mistake)
Looking at your parse method:
for div in divs: l = KikaItemLoader(item=ShowItem(), response=response) # ... add xpaths ... return l.load_item()
Using return inside the loop will exit the function after processing the first <b> element. All other items are never processed, and even the first item might be incomplete if the XPaths fail. Replace return with yield to process all items:
for div in divs: l = KikaItemLoader(item=ShowItem(), response=response) l.add_xpath("title", "./text()") l.add_xpath("date", "./ancestor::ul[1]/preceding-sibling::h2[1]/text()") l.add_xpath("time", "./preceding-sibling::small[1]/text()") yield l.load_item() # Use yield instead of return
2. Fix Allowed Domains
Your allowed_domains is set incorrectly—it should be just the domain, not the full URL:
allowed_domains = ["kinokika.pl"] # Remove http:// and the path
A wrong allowed domain can block your spider from fetching the page properly, leading to empty fields.
3. Validate Your XPath Queries
The most likely culprit is that your XPaths aren't actually pulling the data you expect. Test them in the Scrapy Shell to confirm:
- Open the shell with:
scrapy shell http://www.kinokika.pl/dk.php - Grab the first
<b>element and test each XPath:b = response.xpath('//b')[0] # Test title b.xpath('./text()').get() # Should return the movie title # Test date b.xpath('./ancestor::ul[1]/preceding-sibling::h2[1]/text()').get() # Should return the date # Test time b.xpath('./preceding-sibling::small[1]/text()').get() # Should return the show time
If any of these return None, your XPath is wrong. Inspect the page's HTML structure (right-click > Inspect) to adjust the XPaths—for example, maybe the date is in a different sibling element, or the time uses a <span> instead of <small>.
4. Add Fault Tolerance to the Pipeline
Even if you fix the spider, it's good practice to handle missing fields to avoid crashes. Modify your Pipeline to check for fields before accessing them:
class ScrapyKika(object): def process_item(self, item, spider): # Renamed ShowItem to item (standard practice) # Get or create place first place, created = Place.objects.get_or_create(name="kino kika") # Extract fields safely with .get() show_time = item.get("time") show_date = item.get("date") show_title = item.get("title") # Skip items with missing critical fields if not all([show_time, show_date, show_title]): spider.logger.warning(f"Skipping incomplete item: {item}") return item # Proceed with update_or_create show = Shows.objects.update_or_create( time=show_time, date=show_date, place=place, defaults={'title': show_title} ) return item
Note: I renamed ShowItem to item in the method parameter—using the class name as a variable name is bad practice and can cause confusion.
5. Check Your ItemLoader Configuration
Make sure your strip_string function is working correctly (if it's a custom function). If it's accidentally stripping all content, your fields will be empty. You can replace it with Scrapy's built-in strip if you don't need custom logic:
from scrapy.loader.processors import MapCompose, Join from scrapy.loader import ItemLoader class KikaItemLoader(ItemLoader): title_in = MapCompose(str.strip, str.lower) title_out = Join() time_in = MapCompose(str.strip) time_out = Join() date_in = MapCompose(str.strip) date_out = Join()
Why Your Opera Spider Worked
Your opera spider probably has correct XPaths that reliably fetch all fields, and it uses yield instead of return to process all items. The location difference is a red herring—the issue is entirely in the cinema spider's implementation.
Try these fixes one by one, starting with the loop and allowed domains, then validate your XPaths. That should resolve the KeyError.
内容的提问来源于stack exchange,提问作者Grevioos

