如何获取网站所有文本(含a标签)?三星S9+评测页爬取问题咨询
Got it, let's tackle this problem head-on. The issue you're facing is super common when web scraping—either grabbing too much junk or missing critical content because you're not targeting the right part of the page. Here's the most effective approach for extracting clean article text from that Samsung Galaxy S9+ review (or any similar news/review site):
Step 1: Target the Article's Main Content Container
The root of your problem is searching the entire page for text instead of focusing on the specific DOM element that holds the actual review content. For PhoneArena reviews, the core article text lives inside a dedicated container (you can verify this by inspecting the page with your browser's dev tools).
For that specific S9+ review, the main content is in a <div> with the class review-body. We'll start by isolating this container first.
Step 2: Extract and Clean Text from the Container
Once you have the main container, use BeautifulSoup's stripped_strings generator—it's designed to pull all text nodes from an element (and its children), automatically stripping whitespace from each segment and ignoring empty text nodes. This avoids the mess of findAll(text=True) and the gaps from recursive=False.
Here's a complete code example:
import requests from bs4 import BeautifulSoup # Target URL url = "https://www.phonearena.com/reviews/Samsung-Galaxy-S9-Plus-Review_id4494" # Fetch and parse the page response = requests.get(url) soup = BeautifulSoup(response.text, "html.parser") # Isolate the main article container article_container = soup.find("div", class_="review-body") # Extract and clean the text clean_article_text = "\n\n".join([text for text in article_container.stripped_strings]) # Print or save the result print(clean_article_text)
Step 3: Optional - Remove Unwanted Sub-Elements
If the container still has some junk (like ads, sidebar widgets, or related links nested inside), you can decompose those elements before extracting text:
# Remove unwanted tags from the article container for unwanted_tag in article_container(["script", "style", "aside", "div", "span"]): # Adjust the tag list based on what you find in the page's dev tools unwanted_tag.decompose() # Now extract text as before clean_article_text = "\n\n".join([text for text in article_container.stripped_strings])
Why This Works Better
- No more useless text: By targeting only the review's main container, you skip navigation menus, footers, ads, and other page clutter entirely.
- No missing content: Using
stripped_stringsrecursively pulls text from all child elements inside the container, so you won't miss paragraphs, headings, or bullet points buried in nested tags. - Clean output: The generator handles whitespace automatically, so you don't have to spend extra time cleaning up messy newlines or spaces.
内容的提问来源于stack exchange,提问作者Jomery

