BeautifulSoup提取无标签包裹文本内容的技术求助
Hey Josh, I’ve been in your exact spot when scraping pages where target text sits right in the parent tag alongside nested elements. Here are three simple, reliable ways to grab that attraction name you need:
Method 1: Use .contents to Access the First Text Node
The td tag's .contents property returns a list of all its child nodes—including raw text nodes. Since your attraction name is the first child, you can grab it directly and clean up whitespace with strip():
for td in attraction_names_full: attraction_name = td.contents[0].strip() print(attraction_name) # Output: A Pirate's Adventure - Treasures of the Seven Seas
Method 2: Use .stripped_strings to Pull Ordered Text
BeautifulSoup’s .stripped_strings generator returns all text from the tag and its children, automatically removing extra spaces and newlines. Since your attraction name is the first piece of text, use next() to grab it:
for td in attraction_names_full: attraction_name = next(td.stripped_strings) print(attraction_name)
Method 3: Remove Unwanted Nested Elements First
If you want to clean up the tag structure before extracting text, you can delete the br and span elements using .decompose(), then get the remaining text:
for td in attraction_names_full: # Remove all br and span tags inside the td for child in td.find_all(['br', 'span']): child.decompose() # Extract and clean the leftover text attraction_name = td.get_text(strip=True) print(attraction_name)
All three methods will work perfectly for your scenario—just pick the one that fits your coding style best!
内容的提问来源于stack exchange,提问作者Josh

