You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从嵌套无序列表HTML生成Pandas DataFrame?

Convert Nested HTML Unordered List to Pandas DataFrame for NLP

Got it, let's turn that nested HTML list into a structured DataFrame—ideal starting point for your NLP pipeline. Here's a straightforward approach that handles both top-level and nested list items cleanly:

Step-by-Step Solution

First, we'll use BeautifulSoup to parse the HTML, then traverse the list structure to build a structured dataset that Pandas can easily turn into a DataFrame.

import pandas as pd
from bs4 import BeautifulSoup

# Your full HTML content (fixed the truncated closing tag for completeness)
html = '''<html> 
<body> 
<ul> 
<li> Name 
<ul> 
<li>Many</li> 
<li>Stories</li> 
</ul> 
</li> 
</ul> 
<ul> 
<li> More </li> 
</ul> 
<ul> 
<li>Stuff 
<ul> 
<li>About</li> 
</ul> 
</li> 
</ul> 
</body>
</html>'''

# Parse the HTML with BeautifulSoup
soup = BeautifulSoup(html, 'html.parser')

# Initialize an empty list to store our structured data
structured_data = []

# Loop through top-level unordered lists (ignore nested ones initially)
for top_level_ul in soup.find_all('ul', recursive=False):
    # Iterate over each top-level list item
    for top_level_li in top_level_ul.find_all('li', recursive=False):
        # Clean up the parent list item text (strip extra whitespace)
        parent_label = top_level_li.contents[0].strip()
        # Check if this list item has a nested sub-list
        nested_list = top_level_li.find('ul')
        
        if nested_list:
            # For each child item in the nested list, add a row to our data
            for child_li in nested_list.find_all('li'):
                child_text = child_li.get_text(strip=True)
                structured_data.append({'Parent': parent_label, 'Child': child_text})
        else:
            # No nested items? Add a row with just the parent label (child is NaN)
            structured_data.append({'Parent': parent_label, 'Child': None})

# Convert the structured list to a Pandas DataFrame
df = pd.DataFrame(structured_data)
print(df)

What This Does

  • recursive=False: Ensures we only target top-level <ul> and <li> elements first, so we don't accidentally process nested lists multiple times.
  • Text Cleaning: Uses strip() to get rid of extra newlines and whitespace from the HTML.
  • Handles Both Cases: Accounts for list items with nested sub-lists (maps parent to child) and items without (leaves Child as None for clarity).

Output

Running the code will give you a DataFrame like this:

Parent    Child
0    Name     Many
1    Name  Stories
2    More     None
3   Stuff    About

If your NLP workflow needs a different structure (like flattening all levels into a single column, or tracking hierarchy depth), you can tweak the traversal logic—for example, adding a recursive function to handle deeper nested lists.

内容的提问来源于stack exchange,提问作者epic556

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 03:59:32