如何处理BeautifulSoup爬取字符串的空格去除或左对齐?常规方法失效
Hey there! I see you're struggling with those persistent extra spaces showing up even after using strip() and replace("\n",""). Let's break down why that's happening and fix it right away.
Why Your Current Methods Aren't Working
strip()only removes whitespace from the start and end of the string—it doesn't touch multiple spaces or tabs in the middle.replace("\n","")just gets rid of line breaks, but leaves other whitespace like multiple spaces or tabs intact.ljust()is for aligning text to the left, not cleaning up redundant whitespace.
Two Simple Fixes to Get Clean, Formatted Text
Method 1: Use split() + join() (No Extra Modules Needed)
This is my go-to trick because it's straightforward and doesn't require importing anything new. The split() method automatically splits the string on any sequence of whitespace (spaces, tabs, newlines), and then ' '.join() puts it back together with single spaces between words.
Here's how to update your code:
from bs4 import BeautifulSoup with open('f.html') as f: soup = BeautifulSoup(f, 'lxml') article = soup.findAll('div', class_='modal-content') for i in article: cleaned_text = ' '.join(str(i.text).split()) print(cleaned_text)
Method 2: Regular Expressions (For More Control)
If you want more flexibility (like keeping certain whitespace if needed), you can use regex to replace all sequences of whitespace with a single space. Just import the re module first:
from bs4 import BeautifulSoup import re with open('f.html') as f: soup = BeautifulSoup(f, 'lxml') article = soup.findAll('div', class_='modal-content') for i in article: cleaned_text = re.sub(r'\s+', ' ', str(i.text)).strip() print(cleaned_text)
The regex r'\s+' matches one or more whitespace characters (spaces, tabs, newlines), and replaces them with a single space. The final strip() ensures any leftover leading/trailing space is removed.
Either of these methods should give you the clean, whitespace-free output you're looking for!
内容的提问来源于stack exchange,提问作者Bharath

