You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬取Boardgamegeek:解决点赞数缺失补0问题

Fixing Mismatched List Lengths for 0-Count Recommendations in Scrapy

Great question! The issue here is that when a game has 0 recommendations, the target <a> tag has no text content, so getall() simply skips those entries entirely—leading to your title list being longer than the recommendation count list.

Here's a straightforward fix to ensure both lists stay perfectly aligned, replacing empty values with a default of 0:

Step 1: Fetch all recommendation elements (not just text)

Instead of grabbing text directly with getall(), first retrieve every <a> element that holds the recommendation count—even the empty ones for 0 likes:

# Get all title texts first (this part stays the same)
game_titles = response.css('.fl > a:nth-child(2)::text').getall()

# Fetch ALL recommendation elements, including empty ones for 0 likes
rec_elements = response.css('.recs a.js-score')

Step 2: Process each element to get counts (with 0 fallback)

Loop through each element, extracting its text. If the text is empty (meaning 0 likes), use get(default="0") to set the fallback value:

game_rec_counts = []
for elem in rec_elements:
    # Get text, default to "0" if empty
    raw_count = elem.css('::text').get(default="0")
    # Convert to integer for easier sorting later
    rec_count = int(raw_count) if raw_count.isdigit() else 0
    game_rec_counts.append(rec_count)

Or use a more concise list comprehension if you prefer:

game_rec_counts = [
    int(elem.css('::text').get(default="0")) 
    if elem.css('::text').get(default="0").isdigit() 
    else 0 
    for elem in response.css('.recs a.js-score')
]

Step 3: Pair and sort into your desired dictionary

Now that game_titles and game_rec_counts are the same length, you can zip them into a dictionary and sort by recommendation count:

# Create the initial name:count dictionary
game_rec_dict = dict(zip(game_titles, game_rec_counts))

# Sort the dictionary by count (descending order)
sorted_game_rec_dict = dict(
    sorted(game_rec_dict.items(), key=lambda x: x[1], reverse=True)
)

Why this works

The key difference here is that we're iterating over every individual recommendation element (one per game) instead of relying on getall() to pull text. Since get() supports the default parameter (unlike getall()), we can explicitly set empty entries to 0, ensuring both lists stay in sync.

内容的提问来源于stack exchange,提问作者crystal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 15:07:42