You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python的Beautiful Soup与Selenium从指定URL提取完整表格并导出为与原表格格式一致的CSV文件

Fix for Extracting Japan Meteorological Agency Warning Table to CSV Correctly

I see where your code is going wrong—you're pulling raw text from elements without respecting the table's row-column structure, which is why you're losing empty cells and getting messy column names. Let's rewrite this properly using BeautifulSoup and Selenium to match the original table exactly.

Key Issues in Your Original Code

  • Using time.sleep(3) is unreliable; we need to wait for the table to load explicitly instead of guessing timing
  • Extracting text directly from <tr> tags concatenates all cell content into one string, destroying column separation
  • Empty cells in the original table are being ignored entirely
  • Column naming logic relies on arbitrary text snippets instead of the actual table headers

Corrected Code

from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import pandas as pd

# Configure Chrome options
url = 'https://www.jma.go.jp/bosai/warning/#lang=en'
options = Options()
options.add_argument('--headless')
options.add_argument('--disable-gpu')
options.add_argument('--no-sandbox')  # Optional, fixes environment-specific issues

# Initialize driver (ensure your chromedriver path is correct for your OS)
driver = webdriver.Chrome(options=options)
driver.get(url)

# Wait for the warning table to load (far more reliable than time.sleep)
try:
    # Wait up to 10 seconds for the core warning table to appear
    table = WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.CSS_SELECTOR, "table.warning-table"))
    )
    page_source = driver.page_source
finally:
    driver.quit()

# Parse the page with BeautifulSoup
soup = BeautifulSoup(page_source, 'html.parser')
warning_table = soup.find('table', class_='warning-table')

# Extract table headers
headers = []
for th in warning_table.find('thead').find_all('th'):
    # Clean header text while preserving empty headers if any
    header_text = th.get_text(strip=True)
    headers.append(header_text if header_text else "")

# Extract table rows and cell data
rows_data = []
for tr in warning_table.find('tbody').find_all('tr'):
    row = []
    for td in tr.find_all('td'):
        # Extract cell text, keep empty cells as empty strings
        cell_text = td.get_text(strip=False).strip()
        row.append(cell_text if cell_text else "")
    rows_data.append(row)

# Create DataFrame and save to CSV
df = pd.DataFrame(rows_data, columns=headers)
# Use utf-8-sig to ensure Japanese characters display correctly in Excel
df.to_csv("Japan_Warning_alerts.csv", index=False, encoding='utf-8-sig')

print("CSV saved successfully with exact match to the original table!")

What This Code Does Differently

  • Explicit Loading Wait: Uses WebDriverWait to confirm the table is fully loaded before parsing, no more incomplete data
  • Structured Table Parsing: Iterates through <thead> for headers and <tbody> for rows/cells, preserving the original column order
  • Empty Cell Preservation: Empty cells are kept as empty strings instead of being dropped
  • Clean Text Handling: Balances whitespace stripping while ensuring empty cells aren't mistaken for whitespace-only content
  • Encoding Fix: Uses utf-8-sig to make sure Japanese characters display correctly in spreadsheet tools like Excel

内容的提问来源于stack exchange,提问作者gaurav12

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.27 18:47:51