如何通过Beautiful Soup提取带格式的页面标题与段落?
Got it, let's fix this so you can pull content with proper formatting (like bolded headings) instead of just plain text. The issue with your current code is that it's stripping out all HTML structure and just grabbing raw text. Here's how to adapt it to map HTML tags to Markdown formatting:
First, we'll create a helper function to convert individual HTML elements to their Markdown equivalents. Then we'll iterate through each element inside your target divs and build out formatted content:
from bs4 import BeautifulSoup def html_to_markdown(element): # Handle headings (h1-h6) - apply Markdown heading syntax + bold as requested if element.name in ['h1', 'h2', 'h3', 'h4', 'h5', 'h6']: heading_level = int(element.name[1]) return f"{'#' * heading_level} **{element.get_text(strip=True)}**\n\n" # Handle regular paragraphs elif element.name == 'p': return f"{element.get_text(strip=True)}\n\n" # Handle inline bold text (in case your headings use <strong> tags instead of hX) elif element.name == 'strong': return f"**{element.get_text(strip=True)}**" # Add more tag handlers here (like <em> for italic, <ul>/<li> for lists) as needed else: # Fallback for unknown tags - just extract the text return element.get_text(strip=True) # Your existing setup to target the divs soup = BeautifulSoup(page.content, 'html.parser') target_divs = soup.find_all('div', class_=['x']) formatted_content = [] for div in target_divs: # Iterate through all direct child elements of the div for child in div.children: # Skip empty whitespace nodes that aren't actual HTML tags if child.name is not None: formatted_content.append(html_to_markdown(child)) # Combine all parts into a single Markdown document final_markdown = ''.join(formatted_content) print(final_markdown)
How this works:
- The
html_to_markdownfunction checks each element's tag type and converts it to matching Markdown. For headings, we use the heading level (1-6) to add the right number of#symbols, then wrap the text in**to make it bold as you wanted. - We loop through each direct child of your target
div.xelements, skipping any useless whitespace-only nodes. - Each converted element gets added to a list, which we then join into a clean Markdown string.
You can extend the helper function to handle other tags too—like <em> for italic text (*text*), <ul>/<li> for bullet points, or <a> for links—just add more elif cases tailored to your specific HTML structure.
内容的提问来源于stack exchange,提问作者ctiid

