You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何获取网站所有文本(含a标签)?三星S9+评测页爬取问题咨询

How to Extract Clean Article Text from PhoneArena Reviews

Got it, let's tackle this problem head-on. The issue you're facing is super common when web scraping—either grabbing too much junk or missing critical content because you're not targeting the right part of the page. Here's the most effective approach for extracting clean article text from that Samsung Galaxy S9+ review (or any similar news/review site):

Step 1: Target the Article's Main Content Container

The root of your problem is searching the entire page for text instead of focusing on the specific DOM element that holds the actual review content. For PhoneArena reviews, the core article text lives inside a dedicated container (you can verify this by inspecting the page with your browser's dev tools).

For that specific S9+ review, the main content is in a <div> with the class review-body. We'll start by isolating this container first.

Step 2: Extract and Clean Text from the Container

Once you have the main container, use BeautifulSoup's stripped_strings generator—it's designed to pull all text nodes from an element (and its children), automatically stripping whitespace from each segment and ignoring empty text nodes. This avoids the mess of findAll(text=True) and the gaps from recursive=False.

Here's a complete code example:

import requests
from bs4 import BeautifulSoup

# Target URL
url = "https://www.phonearena.com/reviews/Samsung-Galaxy-S9-Plus-Review_id4494"

# Fetch and parse the page
response = requests.get(url)
soup = BeautifulSoup(response.text, "html.parser")

# Isolate the main article container
article_container = soup.find("div", class_="review-body")

# Extract and clean the text
clean_article_text = "\n\n".join([text for text in article_container.stripped_strings])

# Print or save the result
print(clean_article_text)

Step 3: Optional - Remove Unwanted Sub-Elements

If the container still has some junk (like ads, sidebar widgets, or related links nested inside), you can decompose those elements before extracting text:

# Remove unwanted tags from the article container
for unwanted_tag in article_container(["script", "style", "aside", "div", "span"]):
    # Adjust the tag list based on what you find in the page's dev tools
    unwanted_tag.decompose()

# Now extract text as before
clean_article_text = "\n\n".join([text for text in article_container.stripped_strings])

Why This Works Better

  • No more useless text: By targeting only the review's main container, you skip navigation menus, footers, ads, and other page clutter entirely.
  • No missing content: Using stripped_strings recursively pulls text from all child elements inside the container, so you won't miss paragraphs, headings, or bullet points buried in nested tags.
  • Clean output: The generator handles whitespace automatically, so you don't have to spend extra time cleaning up messy newlines or spaces.

内容的提问来源于stack exchange,提问作者Jomery

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 07:14:01