如何使用BeautifulSoup提取HTML中td标签title属性的结构化数据并转换为DataFrame
Hey there! The issue you're hitting is because tds is a ResultSet (essentially a list) of <td> elements from find_all(), not a single element. You can't access .attrs['title'] directly on the whole list—you need to loop through each individual <td> in that list instead.
Let's walk through a step-by-step solution to get your desired DataFrame:
Step 1: Fix the iteration over <td> elements
Instead of trying to access the list directly, loop through each item in tds. Use .get('title') instead of attrs['title'] to avoid KeyError if a <td> doesn't have a title attribute:
for td in tds: title_text = td.get('title') if not title_text: continue # Skip elements without a title # Process the title string next
Step 2: Parse the title into key-value pairs
Your title format is col1 : val1 col2 : val2—we can use regex to cleanly extract each column and its value. First import the re module, then use re.findall() to grab all pairs:
import re # Inside the loop: col_val_pairs = re.findall(r'(\w+) : (\w+)', title_text) # This returns a list like [('col1', 'val1'), ('col2', 'val2')]
Step 3: Build rows and create the DataFrame
Convert each set of pairs into a dictionary, collect all dictionaries into a list, then pass that list to pandas' DataFrame() constructor:
import pandas as pd data_rows = [] for td in tds: title_text = td.get('title') if not title_text: continue col_val_pairs = re.findall(r'(\w+) : (\w+)', title_text) if col_val_pairs: data_rows.append(dict(col_val_pairs)) # Generate the final DataFrame result_df = pd.DataFrame(data_rows) print(result_df)
Full working code (with your existing setup)
I also fixed a small missing piece in your original code (assuming you meant to select the table from the soup):
from bs4 import BeautifulSoup import re import pandas as pd # Your existing code to fetch and parse the HTML html = driver.page_source soup = BeautifulSoup(html, 'html.parser') table = soup.find('table') # Added: get the table element first tbody = table.select_one('tbody') tds = tbody.find_all("td") # Process tds into DataFrame data_rows = [] for td in tds: title_text = td.get('title') if not title_text: continue col_val_pairs = re.findall(r'(\w+) : (\w+)', title_text) if col_val_pairs: data_rows.append(dict(col_val_pairs)) result_df = pd.DataFrame(data_rows) print(result_df)
This will output exactly the table structure you want, with col1 and col2 as columns and their corresponding values from each <td>'s title attribute.
内容的提问来源于stack exchange,提问作者younghyun

