You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用BeautifulSoup提取HTML中td标签title属性的结构化数据并转换为DataFrame

Hey there! The issue you're hitting is because tds is a ResultSet (essentially a list) of <td> elements from find_all(), not a single element. You can't access .attrs['title'] directly on the whole list—you need to loop through each individual <td> in that list instead.

Let's walk through a step-by-step solution to get your desired DataFrame:

Step 1: Fix the iteration over <td> elements

Instead of trying to access the list directly, loop through each item in tds. Use .get('title') instead of attrs['title'] to avoid KeyError if a <td> doesn't have a title attribute:

for td in tds:
    title_text = td.get('title')
    if not title_text:
        continue  # Skip elements without a title
    # Process the title string next

Step 2: Parse the title into key-value pairs

Your title format is col1 : val1 col2 : val2—we can use regex to cleanly extract each column and its value. First import the re module, then use re.findall() to grab all pairs:

import re

# Inside the loop:
col_val_pairs = re.findall(r'(\w+) : (\w+)', title_text)
# This returns a list like [('col1', 'val1'), ('col2', 'val2')]

Step 3: Build rows and create the DataFrame

Convert each set of pairs into a dictionary, collect all dictionaries into a list, then pass that list to pandas' DataFrame() constructor:

import pandas as pd

data_rows = []
for td in tds:
    title_text = td.get('title')
    if not title_text:
        continue
    col_val_pairs = re.findall(r'(\w+) : (\w+)', title_text)
    if col_val_pairs:
        data_rows.append(dict(col_val_pairs))

# Generate the final DataFrame
result_df = pd.DataFrame(data_rows)
print(result_df)

Full working code (with your existing setup)

I also fixed a small missing piece in your original code (assuming you meant to select the table from the soup):

from bs4 import BeautifulSoup
import re
import pandas as pd

# Your existing code to fetch and parse the HTML
html = driver.page_source
soup = BeautifulSoup(html, 'html.parser')
table = soup.find('table')  # Added: get the table element first
tbody = table.select_one('tbody')
tds = tbody.find_all("td")

# Process tds into DataFrame
data_rows = []
for td in tds:
    title_text = td.get('title')
    if not title_text:
        continue
    col_val_pairs = re.findall(r'(\w+) : (\w+)', title_text)
    if col_val_pairs:
        data_rows.append(dict(col_val_pairs))

result_df = pd.DataFrame(data_rows)
print(result_df)

This will output exactly the table structure you want, with col1 and col2 as columns and their corresponding values from each <td>'s title attribute.

内容的提问来源于stack exchange,提问作者younghyun

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 07:13:09