如何使用Python的Beautiful Soup与Selenium从指定URL提取完整表格并导出为与原表格格式一致的CSV文件
Fix for Extracting Japan Meteorological Agency Warning Table to CSV Correctly
I see where your code is going wrong—you're pulling raw text from elements without respecting the table's row-column structure, which is why you're losing empty cells and getting messy column names. Let's rewrite this properly using BeautifulSoup and Selenium to match the original table exactly.
Key Issues in Your Original Code
- Using
time.sleep(3)is unreliable; we need to wait for the table to load explicitly instead of guessing timing - Extracting text directly from
<tr>tags concatenates all cell content into one string, destroying column separation - Empty cells in the original table are being ignored entirely
- Column naming logic relies on arbitrary text snippets instead of the actual table headers
Corrected Code
from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.chrome.options import Options from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import pandas as pd # Configure Chrome options url = 'https://www.jma.go.jp/bosai/warning/#lang=en' options = Options() options.add_argument('--headless') options.add_argument('--disable-gpu') options.add_argument('--no-sandbox') # Optional, fixes environment-specific issues # Initialize driver (ensure your chromedriver path is correct for your OS) driver = webdriver.Chrome(options=options) driver.get(url) # Wait for the warning table to load (far more reliable than time.sleep) try: # Wait up to 10 seconds for the core warning table to appear table = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CSS_SELECTOR, "table.warning-table")) ) page_source = driver.page_source finally: driver.quit() # Parse the page with BeautifulSoup soup = BeautifulSoup(page_source, 'html.parser') warning_table = soup.find('table', class_='warning-table') # Extract table headers headers = [] for th in warning_table.find('thead').find_all('th'): # Clean header text while preserving empty headers if any header_text = th.get_text(strip=True) headers.append(header_text if header_text else "") # Extract table rows and cell data rows_data = [] for tr in warning_table.find('tbody').find_all('tr'): row = [] for td in tr.find_all('td'): # Extract cell text, keep empty cells as empty strings cell_text = td.get_text(strip=False).strip() row.append(cell_text if cell_text else "") rows_data.append(row) # Create DataFrame and save to CSV df = pd.DataFrame(rows_data, columns=headers) # Use utf-8-sig to ensure Japanese characters display correctly in Excel df.to_csv("Japan_Warning_alerts.csv", index=False, encoding='utf-8-sig') print("CSV saved successfully with exact match to the original table!")
What This Code Does Differently
- Explicit Loading Wait: Uses
WebDriverWaitto confirm the table is fully loaded before parsing, no more incomplete data - Structured Table Parsing: Iterates through
<thead>for headers and<tbody>for rows/cells, preserving the original column order - Empty Cell Preservation: Empty cells are kept as empty strings instead of being dropped
- Clean Text Handling: Balances whitespace stripping while ensuring empty cells aren't mistaken for whitespace-only content
- Encoding Fix: Uses
utf-8-sigto make sure Japanese characters display correctly in spreadsheet tools like Excel
内容的提问来源于stack exchange,提问作者gaurav12
相关产品推荐
相关产品推荐

