Python网页爬取时如何在指定位置为字符串添加空格?
Got it, let's sort out that missing space issue! The problem occurs because when you call .text on the <div class="pin_info">, BeautifulSoup concatenates all text content from child elements directly—ignoring the whitespace that separates those elements in the original HTML. That's why your "Level :2" runs right into "DC Fast :12" without a space.
Simple Fix Using get_text() with Separator
BeautifulSoup has a handy get_text() method that lets you specify a separator for text from different elements. Here's how to use it:
# Get the pin_info div element pin_info_div = browser[0].find("div", {"class": "pin_info"}) # Extract text with spaces between elements, then strip extra whitespace full_clean_text = pin_info_div.get_text(separator=' ', strip=True) # Split into parts and skip the first one (the "A" from the pinID span) result = ' '.join(full_clean_text.split()[1:]) print(result) # Output: "Level :2 DC Fast :12 Level :2"
Alternative: Using stripped_strings
If you prefer more control over individual text segments, use the stripped_strings generator. It returns all non-empty, trimmed text pieces from the element and its children:
pin_info_div = browser[0].find("div", {"class": "pin_info"}) # Get all stripped text parts, skip the first one ("A") text_segments = list(pin_info_div.stripped_strings)[1:] # Join with spaces to get the desired format result = " ".join(text_segments) print(result) # Output: "Level :2 DC Fast :12 Level :2"
Why This Works
- Both methods ensure that text from separate child elements (like the
<span class="dark">and others) are separated by spaces. - They automatically handle any extra whitespace or empty text nodes (like from the
<i>tag which has no text content). - Skipping the first segment removes the "A" from the
<span class="pinID">which you were already doing with[1:]in your original code.
Either approach should give you the formatted output you're looking for!
内容的提问来源于stack exchange,提问作者Eunice TT

