You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过Beautiful Soup提取带格式的页面标题与段落?

Got it, let's fix this so you can pull content with proper formatting (like bolded headings) instead of just plain text. The issue with your current code is that it's stripping out all HTML structure and just grabbing raw text. Here's how to adapt it to map HTML tags to Markdown formatting:

First, we'll create a helper function to convert individual HTML elements to their Markdown equivalents. Then we'll iterate through each element inside your target divs and build out formatted content:

from bs4 import BeautifulSoup

def html_to_markdown(element):
    # Handle headings (h1-h6) - apply Markdown heading syntax + bold as requested
    if element.name in ['h1', 'h2', 'h3', 'h4', 'h5', 'h6']:
        heading_level = int(element.name[1])
        return f"{'#' * heading_level} **{element.get_text(strip=True)}**\n\n"
    # Handle regular paragraphs
    elif element.name == 'p':
        return f"{element.get_text(strip=True)}\n\n"
    # Handle inline bold text (in case your headings use <strong> tags instead of hX)
    elif element.name == 'strong':
        return f"**{element.get_text(strip=True)}**"
    # Add more tag handlers here (like <em> for italic, <ul>/<li> for lists) as needed
    else:
        # Fallback for unknown tags - just extract the text
        return element.get_text(strip=True)

# Your existing setup to target the divs
soup = BeautifulSoup(page.content, 'html.parser')
target_divs = soup.find_all('div', class_=['x'])
formatted_content = []

for div in target_divs:
    # Iterate through all direct child elements of the div
    for child in div.children:
        # Skip empty whitespace nodes that aren't actual HTML tags
        if child.name is not None:
            formatted_content.append(html_to_markdown(child))

# Combine all parts into a single Markdown document
final_markdown = ''.join(formatted_content)
print(final_markdown)

How this works:

  1. The html_to_markdown function checks each element's tag type and converts it to matching Markdown. For headings, we use the heading level (1-6) to add the right number of # symbols, then wrap the text in ** to make it bold as you wanted.
  2. We loop through each direct child of your target div.x elements, skipping any useless whitespace-only nodes.
  3. Each converted element gets added to a list, which we then join into a clean Markdown string.

You can extend the helper function to handle other tags too—like <em> for italic text (*text*), <ul>/<li> for bullet points, or <a> for links—just add more elif cases tailored to your specific HTML structure.

内容的提问来源于stack exchange,提问作者ctiid

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 03:55:26