You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提取CSS类文本并将网页中公司详情转换为表格?

Extract Company Details into a Markdown Table with BeautifulSoup

Looks like you're halfway there with your code—let's fix and expand it to properly pull those address components and turn them into a clean table. Here's a complete, working solution tailored to your needs:

Full Working Code

from urllib.request import urlopen
from bs4 import BeautifulSoup
import pandas as pd

# Fetch and parse the target HTML
html_doc = 'https://s3.amazonaws.com/todel162/test.html'
soup = BeautifulSoup(urlopen(html_doc), 'html.parser')

# Store extracted details in a list of dictionaries
company_info = []

# Loop through each row containing address fields
for row in soup.find_all("div", class_="row"):
    # Grab the columns (replace the class selector with the full one from your HTML)
    columns = row.find_all("div", class_="col-md-3 col-sm-6")  # Use the exact class name here
    # Make sure we have both label and value columns
    if len(columns) == 2:
        label = columns[0].get_text(strip=True)
        value = columns[1].get_text(strip=True)
        company_info.append({"Detail": label, "Value": value})

# Option 1: Use pandas to generate a markdown table quickly
if company_info:
    df = pd.DataFrame(company_info)
    print(df.to_markdown(index=False))
else:
    print("No company details found.")

# Option 2: Manual markdown table (no pandas required)
def create_markdown_table(data):
    if not data:
        return "No data extracted."
    # Build table headers
    headers = list(data[0].keys())
    table = f"| {' | '.join(headers)} |\n| {' | '.join(['---']*len(headers))} |\n"
    # Add rows
    for entry in data:
        table += f"| {' | '.join(entry.values())} |\n"
    return table

print(create_markdown_table(company_info))

Key Fixes & Explanations

  • Truncated Class Selector: Your original code had col-md-3 col...—you need to use the full, exact class name from the HTML (inspect elements with F12 to get this right). The class_ parameter in BeautifulSoup is cleaner than the old dictionary syntax.
  • Clean Text Extraction: get_text(strip=True) strips out messy whitespace and newlines, so your extracted values are neat.
  • Table Options:
    • Pandas is the fastest way to turn your data into a markdown table—just install it with pip install pandas if you don't have it.
    • The manual function works if you want to avoid dependencies, building a proper markdown table structure from scratch.

Example Output

If your HTML has the address fields you mentioned, the table will look like this:

DetailValue
Block NumberXXX
Building NameYYY
Street Namezzz
Postal Code123456789

Quick Troubleshooting

  • Double-check Class Names: If you're getting no results, use your browser's dev tools to confirm the exact class names of the divs holding the labels and values.
  • Handle Missing Entries: Add checks to skip rows where columns are missing to avoid index errors.
  • Multiple Companies: If the page has multiple companies, wrap the row loop in another loop that targets each company's container div first.

内容的提问来源于stack exchange,提问作者shantanuo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 08:01:11