You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页<div>内表格数据提取及Excel单元格映射问题求助

Fixing Table Extraction to Excel: Avoiding Single-Cell Data Dump

Hey Francis, let's tackle this table extraction issue you're having—sounds like you've already nailed the hard parts (login, navigation, loading the report!), so we just need to get the data mapping right for Excel. The problem where all your table content ends up in one cell usually happens when you're not properly splitting the table into rows and individual cells before writing to Excel. Let's walk through solutions for the two most common scraping tools: BeautifulSoup (for static HTML) and Selenium (for dynamic content).

1. If You're Using BeautifulSoup (Static HTML)

First, we'll parse the table into a 2D list (each sublist represents a row, each element a cell), then write that list to Excel row-by-row, cell-by-cell.

Example Code:

from openpyxl import Workbook
from bs4 import BeautifulSoup

# Assume you already have your page's HTML content stored here
page_html = # Your loaded page HTML (after login/navigation)

# Step 1: Locate the table inside your target div
soup = BeautifulSoup(page_html, "html.parser")
target_div = soup.find("div", id="your-div-id")  # Use class_ instead of id if needed
table = target_div.find("table")

# Step 2: Extract table data into a 2D list
table_data = []
for row in table.find_all("tr"):
    # Grab both header cells (<th>) and data cells (<td>)
    cells = row.find_all(["th", "td"])
    # Clean up text (remove extra newlines/spaces) and build row data
    row_content = [cell.get_text(strip=True) for cell in cells]
    table_data.append(row_content)

# Step 3: Write to Excel correctly
wb = Workbook()
ws = wb.active

# Loop through the 2D list and map each value to its own cell
for row_num, row in enumerate(table_data, start=1):
    for col_num, cell_value in enumerate(row, start=1):
        ws.cell(row=row_num, column=col_num, value=cell_value)

wb.save("extracted_table.xlsx")

2. If You're Using Selenium (Dynamic Content)

Since you mentioned loading the report, if the table is rendered dynamically (JavaScript), Selenium is likely your tool of choice. The approach is similar—extract rows and cells first, then map to Excel.

Example Code:

from selenium import webdriver
from selenium.webdriver.common.by import By
from openpyxl import Workbook

# Initialize driver and complete login/navigation (you already have this part!)
driver = webdriver.Chrome()
# ... your login/navigation code here ...

# Step 1: Locate the table inside the target div
target_table = driver.find_element(By.CSS_SELECTOR, "div.your-div-class table")  # Adjust selector
rows = target_table.find_elements(By.TAG_NAME, "tr")

# Step 2: Extract table data into a 2D list
table_data = []
for row in rows:
    # Get all cells in the row (headers + data)
    cells = row.find_elements(By.TAG_NAME, ["td", "th"])
    row_content = [cell.text.strip() for cell in cells]
    table_data.append(row_content)

# Step 3: Write to Excel (same as BeautifulSoup example)
wb = Workbook()
ws = wb.active

for row_num, row in enumerate(table_data, start=1):
    for col_num, cell_value in enumerate(row, start=1):
        ws.cell(row=row_num, column=col_num, value=cell_value)

wb.save("selenium_table.xlsx")
driver.quit()

Why Your Previous Code Failed

Chances are you were doing one of these:

  • Grabbing the entire table's raw text (e.g., table.get_text() or table.text) and dumping that single string into one Excel cell.
  • Not iterating through each row and cell individually, so you never split the data into discrete values.

Quick Debug Tip

Before writing to Excel, print out table_data—you should see a list of lists, like:

[["Header 1", "Header 2"], ["Row 1 Col1", "Row1 Col2"], ["Row2 Col1", "Row2 Col2"]]

If that's what you see, the Excel write will work correctly. If you see a single long string, go back to how you're extracting rows/cells.

内容的提问来源于stack exchange,提问作者Francis

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 04:18:59