You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

BeautifulSoup爬取Steam文明6评论导出CSV格式异常问题求助

代码存在的两个核心问题

  • 缩进错误导致仅存单条数据
    你将review_dict的列表追加逻辑放在了for cards in card_div循环的外部,循环遍历所有评论卡片后,只会把最后一次循环赋值的变量追加到列表中,最终只会保留1条数据。需要将所有append语句缩进一层,放入循环内部,保证每条评论遍历后都被存入列表。
  • 未提取纯文本导致HTML标签残留
    find()方法返回的是BeautifulSoup的Tag对象,直接存入列表导出CSV时会连带标签结构一起输出。你需要对每个提取到的Tag对象调用.text.strip()方法获取纯文本内容,空值场景要做判断避免报错;另外评论文本长度不能直接对Tag对象取len,要先拿到纯文本再统计长度。

修正后的可运行代码

import pandas as pd
import requests
from bs4 import BeautifulSoup as bs

# 加请求头避免Steam反爬拦截
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36"
}
url = "https://steamcommunity.com/app/289070/reviews/?browsefilter=toprated&snr=1_5_100010_"

review_dict = {
    "found_helpful": [],
    "title": [],  # 推荐或不推荐
    "hours": [],
    "prods_in_account": [],
    "words_in_review": []
}

def data_scrapper():
    """
    从Steam页面获取评论数据
    """
    response = requests.get(url, headers=headers)
    soup = bs(response.content, "html.parser")
    card_div = soup.find_all("div", attrs={"class": "apphub_Card modalContentLink interactable"})

    for cards in card_div:
        # 提取有用次数,过滤多余空白字符
        found_helpful = cards.find("div", attrs={"class": "found_helpful"})
        found_helpful_text = found_helpful.text.strip() if found_helpful else ""
        
        # 提取推荐状态
        vote_header = cards.find("div", attrs={"class": "title"}) # 推荐状态直接在.title标签下
        vote_header_text = vote_header.text.strip() if vote_header else ""
        
        # 提取游戏时长
        hours = cards.find("div", attrs={"class": "hours"})
        hours_text = hours.text.strip() if hours else ""
        
        # 提取账号持有产品数
        products = cards.find("div", attrs={"class": "apphub_CardContentMoreLink ellipsis"})
        products_text = products.text.strip() if products else ""
        
        # 提取评论文本并统计长度
        words_in_review = cards.find("div", attrs={"class": "apphub_CardTextContent"})
        review_text = words_in_review.text.strip() if words_in_review else ""
        review_length = len(review_text)

        # 缩进放在循环内部,每条评论都追加
        review_dict["found_helpful"].append(found_helpful_text)
        review_dict["title"].append(vote_header_text)
        review_dict["hours"].append(hours_text)
        review_dict["prods_in_account"].append(products_text)
        review_dict["words_in_review"].append(review_length)

data_scrapper()

review_df = pd.DataFrame.from_dict(review_dict)
review_df.to_csv("review.csv", sep=",", encoding="utf_8_sig", index=False)

补充说明

修正后代码新增了请求头避免Steam反爬拦截,导出CSV时指定了utf_8_sig编码避免中文乱码,同时新增了空值判断逻辑,避免页面元素缺失时代码报错。


内容的提问来源于stack exchange,提问作者Karthik

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.29 23:57:03