为何stripped_string输出无逗号分隔?如何提取每个<li>最后文本?
Fixing Your BeautifulSoup Text Extraction
Hey there! Let's tweak your code to get exactly the output you're after—grabbing only the last text segment from each li.odd element, formatted as single-item lists.
The Core Issue & Fix
Your current code captures all text segments split by <br> tags (via stripped_strings), but you only need the last one. Here's how to adjust it:
- Grab the last element from the stripped strings list: When you convert
li.stripped_stringsto a list, using the index[-1]will always give you the final text segment (which corresponds to the content after the last<br>). - Handle empty cases (optional but safe): Add a check to avoid index errors if any
lihas no text content. - Format output as single-item lists: Wrap the last text segment in square brackets when printing.
Modified Working Code
from bs4 import BeautifulSoup import requests URL = '...' content = requests.get(URL) soup = BeautifulSoup(content.text, 'lxml') contentTable = soup.find_all("li", {"class": "odd"}) for li in contentTable: # Convert the stripped strings iterator to a list text_segments = list(li.stripped_strings) # Only proceed if there's content to avoid index errors if text_segments: # Get the last segment (after the final <br>) target_text = text_segments[-1] # Print as a single-item list print([target_text])
Why This Works
stripped_stringsautomatically ignores<br>tags and extracts all text chunks, stripping extra whitespace from each.- Using
[-1]on the list is a Python shortcut to access the last element, regardless of how many<br>tags are present in theli. - The empty list check ensures your code doesn't crash if an
li.oddelement has no text content.
Expected Output
Running this code will give you exactly what you wanted:
['text4'] ['text7'] ['text11'] ...
内容的提问来源于stack exchange,提问作者Keyvan Darnell
相关产品推荐
相关产品推荐

