You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python BeautifulSoup提取Floor Activity至Vote(s)表格区间的HTML内容

Solution to Extract HTML Tables Between "Floor Activity" and "Vote(s)" Sections

Got it, let's fix this step by step. You're trying to extract all tables between the "Floor Activity" section and the "Vote(s)" section, then package that into a properly escaped JSON. Here's how to do it with BeautifulSoup correctly, avoiding the NoneType error you hit earlier:

Step 1: Locate the Starting Point ("Floor Activity" Table)

First, we need to accurately find the table that marks the start of your target content. Instead of relying on brittle parent chains or index-based XPath, we'll trace up from the specific font tag with "Floor Activity" text:

from bs4 import BeautifulSoup, NavigableString
import json

# Assume you've already fetched the HTML content (e.g., via requests.get(url).text)
html_content = ... # Insert your full HTML content here
soup = BeautifulSoup(html_content, 'html.parser')

# Find the font tag containing "Floor Activity" (handles whitespace around text)
fa_font = soup.find('font', string=lambda text: text and 'Floor Activity' in text.strip())
if not fa_font:
    raise ValueError("Could not locate the 'Floor Activity' section")

# Trace up to the top-level table that contains this section header
fa_table = fa_font.find_parent('table')
while fa_table.find_parent('table'):
    fa_table = fa_table.find_parent('table')

Step 2: Collect Tables Until "Vote(s)" Section

The NoneType error happened because next_sibling often returns non-tag nodes (like newlines or whitespace). We'll filter these out and stop as soon as we hit the "Vote(s)" section:

target_tables = []
current_sibling = fa_table.next_sibling

while current_sibling:
    # Skip non-tag nodes (whitespace, newlines, etc.)
    if isinstance(current_sibling, NavigableString):
        current_sibling = current_sibling.next_sibling
        continue
    
    # Check if we've reached the "Vote(s)" section
    vote_font = current_sibling.find('font', string=lambda text: text and 'Vote(s)' in text.strip())
    if vote_font:
        break
    
    # Add valid tables to our collection
    if current_sibling.name == 'table':
        target_tables.append(str(current_sibling))
    
    current_sibling = current_sibling.next_sibling

Step 3: Generate the Escaped JSON Output

We'll combine the collected tables into a single string, then use json.dumps to automatically handle all special character escaping (quotes, newlines, etc.):

# Combine all table HTML into one string
flooract_html = '\n'.join(target_tables)

# Build the required JSON structure
output_json = {
    "title": "Floor Activity",
    "body": flooract_html
}

# Convert to properly escaped JSON (handles all special characters)
escaped_json = json.dumps(output_json, indent=2, ensure_ascii=False)

# Write to a file if needed
with open('floor_activity_output.json', 'w', encoding='utf-8') as f:
    f.write(escaped_json)

Why This Works

  • Avoids NoneType Errors: We explicitly skip non-tag sibling nodes, so we only interact with valid HTML elements.
  • Reliable Section Boundaries: String matching on the font tags ensures we don't break if the page's table count or structure changes.
  • Valid JSON Output: json.dumps takes care of escaping all HTML special characters, so your final JSON is compliant and ready to use.

Example Output Snippet

The resulting JSON will match your requested format, with an escaped HTML body:

{
  "title": "Floor Activity",
  "body": "<table>\n<tbody>\n<tr>\n<td class=\"labelplain\">&nbsp;Status Date<\/td><td class=\"labelplain\">&nbsp;10/12/2005<\/td>\n<\/tr>\n<tr><td class=\"labelplain\">&nbsp;Parliamentary Status<\/td>\n<td class=\"labelplain\"><table>\n<tbody><tr>\n<td class=\"labelplain\">SPONSORSHIP SPEECH<br>...Until Period of Committee Amendments<\/td>\n<\/tr>\n<\/tbody>\n<\/table><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n..."
}

内容的提问来源于stack exchange,提问作者schnydszch

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 16:37:33