You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup提取HTML表格列并存入Python DataFrame?

Extracting Column 1 Data into a Pandas DataFrame

Got it, let's walk through how to pull that Column 1 content and get it into a pandas DataFrame. I'll build off the code you already have, with step-by-step explanations:

Step 1: Add Pandas Import

First, you'll need pandas to create the DataFrame—don't forget to import it:

import pandas as pd

Step 2: Locate the Target Table

HTML tables live in <table> tags. If your page only has one table, use soup.find('table'). If there are multiple tables, narrow it down with a specific class or ID (check the page's HTML source for this):

from bs4 import BeautifulSoup as BS
import urllib.request
import pandas as pd

url = "your-target-url-here"  # Replace with your actual URL
html = urllib.request.urlopen(url)
soup = BS(html, "lxml")

# Grab the table—adjust the selector if needed (e.g., by class)
table = soup.find('table')  # Or soup.find('table', class_='your-table-class')

Step 3: Extract Column 1 Content

Now we'll loop through each row in the table, pull the first cell (which is Column 1), and collect the text into a list. We'll skip empty rows or header rows if needed:

column1_data = []

# Loop through each row in the table
for row in table.find_all('tr'):
    # Get all cells in the row
    cells = row.find_all('td')
    # Only process rows that have cells (skip header rows or empty rows)
    if cells:
        # Extract text from the first cell (Column 1), strip extra whitespace
        column1_text = cells[0].get_text(strip=True)
        column1_data.append(column1_text)

Alternative: Target Column by Header Text

If you want to be more precise (in case Column 1 isn't always the first cell), you can first map the header names to their positions:

# Extract header names
headers = [header.get_text(strip=True) for header in table.find_all('th')]
# Find the index of "Column 1" in the headers
column1_idx = headers.index('Column 1')

# Now extract the correct column from each row
column1_data = []
for row in table.find_all('tr')[1:]:  # Skip the header row
    cells = row.find_all('td')
    if cells:
        column1_text = cells[column1_idx].get_text(strip=True)
        column1_data.append(column1_text)

Step 4: Create the DataFrame

Finally, turn your list of Column 1 data into a pandas DataFrame:

df = pd.DataFrame({'Column 1': column1_data})
# Check the result
print(df)

Quick Notes for Edge Cases

  • If the table uses <thead> and <tbody> tags, target the body rows with table.find('tbody').find_all('tr') instead.
  • If cells contain links or other elements, adjust the extraction: use cells[0].find('a')['href'] to get a link URL, or cells[0].get_text() to pull all text inside the cell.

内容的提问来源于stack exchange,提问作者Skyhopper9

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 12:29:24