BeautifulSoup爬取Steam文明6评论导出CSV格式异常问题求助
代码存在的两个核心问题
- 缩进错误导致仅存单条数据
你将review_dict的列表追加逻辑放在了for cards in card_div循环的外部,循环遍历所有评论卡片后,只会把最后一次循环赋值的变量追加到列表中,最终只会保留1条数据。需要将所有append语句缩进一层,放入循环内部,保证每条评论遍历后都被存入列表。 - 未提取纯文本导致HTML标签残留
find()方法返回的是BeautifulSoup的Tag对象,直接存入列表导出CSV时会连带标签结构一起输出。你需要对每个提取到的Tag对象调用.text.strip()方法获取纯文本内容,空值场景要做判断避免报错;另外评论文本长度不能直接对Tag对象取len,要先拿到纯文本再统计长度。
修正后的可运行代码
import pandas as pd import requests from bs4 import BeautifulSoup as bs # 加请求头避免Steam反爬拦截 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36" } url = "https://steamcommunity.com/app/289070/reviews/?browsefilter=toprated&snr=1_5_100010_" review_dict = { "found_helpful": [], "title": [], # 推荐或不推荐 "hours": [], "prods_in_account": [], "words_in_review": [] } def data_scrapper(): """ 从Steam页面获取评论数据 """ response = requests.get(url, headers=headers) soup = bs(response.content, "html.parser") card_div = soup.find_all("div", attrs={"class": "apphub_Card modalContentLink interactable"}) for cards in card_div: # 提取有用次数,过滤多余空白字符 found_helpful = cards.find("div", attrs={"class": "found_helpful"}) found_helpful_text = found_helpful.text.strip() if found_helpful else "" # 提取推荐状态 vote_header = cards.find("div", attrs={"class": "title"}) # 推荐状态直接在.title标签下 vote_header_text = vote_header.text.strip() if vote_header else "" # 提取游戏时长 hours = cards.find("div", attrs={"class": "hours"}) hours_text = hours.text.strip() if hours else "" # 提取账号持有产品数 products = cards.find("div", attrs={"class": "apphub_CardContentMoreLink ellipsis"}) products_text = products.text.strip() if products else "" # 提取评论文本并统计长度 words_in_review = cards.find("div", attrs={"class": "apphub_CardTextContent"}) review_text = words_in_review.text.strip() if words_in_review else "" review_length = len(review_text) # 缩进放在循环内部,每条评论都追加 review_dict["found_helpful"].append(found_helpful_text) review_dict["title"].append(vote_header_text) review_dict["hours"].append(hours_text) review_dict["prods_in_account"].append(products_text) review_dict["words_in_review"].append(review_length) data_scrapper() review_df = pd.DataFrame.from_dict(review_dict) review_df.to_csv("review.csv", sep=",", encoding="utf_8_sig", index=False)
补充说明
修正后代码新增了请求头避免Steam反爬拦截,导出CSV时指定了utf_8_sig编码避免中文乱码,同时新增了空值判断逻辑,避免页面元素缺失时代码报错。
内容的提问来源于stack exchange,提问作者Karthik
相关产品推荐
相关产品推荐

