Python+Beautiful Soup爬取Steam数据:平台转0/1值异常求助
Hey there, let's figure out why all your platform values are showing "1"! The main issues are in how you're initializing and checking the platform support status. Let's break down the problems and fix your code step by step.
What's Wrong with Your Current Code?
- Uninitialized/Unreset Variables: You never set default values for platform support (like "0") at the start of each product loop. If a game doesn't trigger your platform check loop, the variable keeps its value from the previous game.
- Overcomplicated Platform Checks: Using regex loops for each platform is unnecessary, and if a platform isn't present for a game, your loop never runs—so the value stays stuck as the last one.
- Undefined Variable: The
yvariable in your final CSV write line isn't defined anywhere, which would throw an error. - No Missing Data Handling: Your code assumes every element (like release date, price) exists, which can crash if a game has incomplete info.
Fixed Code
from urllib.request import urlopen as uReq from bs4 import BeautifulSoup as soup import re from datetime import datetime import time import csv my_url = 'https://store.steampowered.com/search/?specials=1&page=1' # Opening connection, grabbing the page uClient = uReq(my_url) page_html = uClient.read() uClient.close() # HTML parsing page_soup = soup(page_html, "html.parser") # Grab products containers = page_soup.findAll("div", {"class":"responsive_search_name_combined"}) filename = "products.csv" # Add newline='' to avoid extra blank lines in CSV f = open(filename, "w", encoding='UTF-8', newline='') headers = "Titles, Release_date, Discount, Price before, Price after, Positive review, Reviewers, Win, Lin, Osx, Time\n" f.write(headers) ts = time.time() st = datetime.fromtimestamp(ts).strftime('%Y-%m-%d %H:%M:%S') print(st) for container in containers: # Initialize platform support to 0 by default (critical fix!) win_support = "0" lin_support = "0" osx_support = "0" # Get title titles_container = container.findAll("span",{"class":"title"}) titl = titles_container[0].text.strip() print(titl) # Get release date (handle missing dates) product_container = container.findAll("div",{"class":"search_released"}) product_date = product_container[0].text.strip() if product_container else "N/A" print(product_date) # Get discount (handle no discount cases) product_discount_container = container.findAll("div",{"class":"search_discount"}) product_discount = product_discount_container[0].text.strip() if product_discount_container else "0%" print(product_discount) # Get original price (better regex for decimals) product_price_container_before = container.findAll("div",{"class":"search_price"}) product_price_before = product_price_container_before[0].text.strip() if product_price_container_before else "N/A" test = re.findall(r'(\d+\.?\d*)', product_price_before) testing = test[0] if len(test) >=1 else "N/A" print(testing) # Get discounted price (clean up messy whitespace) product_price_after = "N/A" product_price_container_after = container.findAll("div",{"class":"discounted"}) for elem in product_price_container_after: for span in elem.findAll("span"): span.extract() product_price_after = elem.text.strip().replace("\t", "").replace(" ", "") print(product_price_after) # Get review data (handle missing reviews) positive_review = "N/A" reviewers = "N/A" product_review_container = container.findAll("span",{"class":"search_review_summary"}) for product_review in product_review_container: prr = product_review.get('data-tooltip-html') if prr: matches = re.findall(r'(\d+)%.*?(\d+) user', prr) if matches: positive_review = f"{matches[0][0]}%" reviewers = matches[0][1] print(f"Positive: {positive_review}, Reviewers: {reviewers}") # Simplified platform check (directly match class names) product_platform_container = container.findAll("span",{"class":"platform_img"}) for platform in product_platform_container: platform_classes = platform.get('class') if len(platform_classes) >=2: platform_type = platform_classes[1] if platform_type == "win": win_support = "1" elif platform_type == "mac": osx_support = "1" elif platform_type == "linux": lin_support = "1" print(f"Win: {win_support}, Lin: {lin_support}, OSX: {osx_support}") # Write to CSV safely (escape commas in titles) titl_escaped = titl.replace('"', '""') # Escape double quotes for valid CSV f.write(f'"{titl_escaped}",{product_date.replace(",","")},"{product_discount}",{testing},{product_price_after},"{positive_review}",{reviewers},{win_support},{lin_support},{osx_support},"{st}"\n') f.close() # Optional: Verify the CSV import pandas as pd df = pd.read_csv("products.csv") print(df.head())
Key Fixes & Improvements
- Default Platform Values: We start each game loop with all platforms set to "0"—this ensures we don't carry over values from previous games.
- Direct Platform Checks: Instead of regex, we look at the exact class names (
win,mac,linux) and flip the corresponding value to "1" if present. - Missing Data Handling: Added checks for cases where elements like release dates or reviews are missing, so your code won't crash unexpectedly.
- CSV Safety: Escaped commas in game titles and used proper quoting to avoid breaking the CSV structure.
- Cleaner Regex: Updated regex for prices and reviews to handle different formats more reliably.
- Defined Variables: Replaced the undefined
ywithwin_supportand properly addedlin_supportto match your CSV headers.
内容的提问来源于stack exchange,提问作者Sofelia
相关产品推荐
相关产品推荐

