使用Python BeautifulSoup提取Floor Activity至Vote(s)表格区间的HTML内容
Got it, let's fix this step by step. You're trying to extract all tables between the "Floor Activity" section and the "Vote(s)" section, then package that into a properly escaped JSON. Here's how to do it with BeautifulSoup correctly, avoiding the NoneType error you hit earlier:
Step 1: Locate the Starting Point ("Floor Activity" Table)
First, we need to accurately find the table that marks the start of your target content. Instead of relying on brittle parent chains or index-based XPath, we'll trace up from the specific font tag with "Floor Activity" text:
from bs4 import BeautifulSoup, NavigableString import json # Assume you've already fetched the HTML content (e.g., via requests.get(url).text) html_content = ... # Insert your full HTML content here soup = BeautifulSoup(html_content, 'html.parser') # Find the font tag containing "Floor Activity" (handles whitespace around text) fa_font = soup.find('font', string=lambda text: text and 'Floor Activity' in text.strip()) if not fa_font: raise ValueError("Could not locate the 'Floor Activity' section") # Trace up to the top-level table that contains this section header fa_table = fa_font.find_parent('table') while fa_table.find_parent('table'): fa_table = fa_table.find_parent('table')
Step 2: Collect Tables Until "Vote(s)" Section
The NoneType error happened because next_sibling often returns non-tag nodes (like newlines or whitespace). We'll filter these out and stop as soon as we hit the "Vote(s)" section:
target_tables = [] current_sibling = fa_table.next_sibling while current_sibling: # Skip non-tag nodes (whitespace, newlines, etc.) if isinstance(current_sibling, NavigableString): current_sibling = current_sibling.next_sibling continue # Check if we've reached the "Vote(s)" section vote_font = current_sibling.find('font', string=lambda text: text and 'Vote(s)' in text.strip()) if vote_font: break # Add valid tables to our collection if current_sibling.name == 'table': target_tables.append(str(current_sibling)) current_sibling = current_sibling.next_sibling
Step 3: Generate the Escaped JSON Output
We'll combine the collected tables into a single string, then use json.dumps to automatically handle all special character escaping (quotes, newlines, etc.):
# Combine all table HTML into one string flooract_html = '\n'.join(target_tables) # Build the required JSON structure output_json = { "title": "Floor Activity", "body": flooract_html } # Convert to properly escaped JSON (handles all special characters) escaped_json = json.dumps(output_json, indent=2, ensure_ascii=False) # Write to a file if needed with open('floor_activity_output.json', 'w', encoding='utf-8') as f: f.write(escaped_json)
Why This Works
- Avoids
NoneTypeErrors: We explicitly skip non-tag sibling nodes, so we only interact with valid HTML elements. - Reliable Section Boundaries: String matching on the
fonttags ensures we don't break if the page's table count or structure changes. - Valid JSON Output:
json.dumpstakes care of escaping all HTML special characters, so your final JSON is compliant and ready to use.
Example Output Snippet
The resulting JSON will match your requested format, with an escaped HTML body:
{ "title": "Floor Activity", "body": "<table>\n<tbody>\n<tr>\n<td class=\"labelplain\"> Status Date<\/td><td class=\"labelplain\"> 10/12/2005<\/td>\n<\/tr>\n<tr><td class=\"labelplain\"> Parliamentary Status<\/td>\n<td class=\"labelplain\"><table>\n<tbody><tr>\n<td class=\"labelplain\">SPONSORSHIP SPEECH<br>...Until Period of Committee Amendments<\/td>\n<\/tr>\n<\/tbody>\n<\/table><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n..." }
内容的提问来源于stack exchange,提问作者schnydszch

