如何基于BeautifulSoup4实现高效编码?——parse_table_data函数性能优化问询
Optimizing Your BeautifulSoup4 Code for Speed
First, let's look at how to speed up your existing BeautifulSoup implementation without switching libraries—these changes can cut down on unnecessary DOM traversals and redundant checks:
Use a Faster Parser
The built-inhtml.parseris reliable but slow. Switching tolxml(a C-based parser) will immediately speed up the initial parsing step. You'll need to install it first withpip install lxml.Combine Selectors to Reduce Traversals
Instead of finding allscaledRoad--7fdfbdivs first and then looping through each to find SVGs, use a combined CSS selector to target only the relevant SVGs directly. This minimizes the number of times you traverse the DOM.Optimize Condition Checks
Replacing substring checks ("Banker" in name) with set membership (first_part in {'Banker', 'Player'}) is faster because set lookups are O(1). Also, splitting the name once and checking the first part avoids redundant string operations.
Here's the optimized BeautifulSoup code:
import typing from bs4 import BeautifulSoup def parse_table_data(self) -> typing.Union[dict, None]: page_source = self.driver.page_source # Use lxml for faster parsing soup = BeautifulSoup(page_source, "lxml") road_result_container = {"A": [], "B": [], "C": [], "D": [], "E": [], "F": []} column_letters = ["A", "B", "C", "D", "E", "F"] # Get all column divs in order (assuming they map to A-F) column_divs = soup.select("div.scaledRoad--7fdfb") for idx, div in enumerate(column_divs): if idx >= len(column_letters): break # Handle unexpected extra divs gracefully current_column = column_letters[idx] # Target only relevant SVGs within this column for svg in div.select('svg.svg--34293[name][data-type="roadItem"]'): name = svg["name"] first_part = name.split()[0] if first_part in {"Banker", "Player"}: road_result_container[current_column].append(first_part) return road_result_container
Alternative Libraries for Maximum Speed
If you still need more performance after optimizing BeautifulSoup, consider using lxml directly (it's the engine behind BeautifulSoup's fast mode) or pyquery (which uses jQuery-style selectors). Both avoid the abstraction layer of BeautifulSoup, leading to faster DOM operations.
Example with lxml Directly:
import typing from lxml import html def parse_table_data(self) -> typing.Union[dict, None]: page_source = self.driver.page_source tree = html.fromstring(page_source) road_result_container = {"A": [], "B": [], "C": [], "D": [], "E": [], "F": []} column_letters = ["A", "B", "C", "D", "E", "F"] # XPath to get all scaledRoad divs column_divs = tree.xpath('//div[contains(@class, "scaledRoad--7fdfb")]') for idx, div in enumerate(column_divs): if idx >= len(column_letters): break current_column = column_letters[idx] # XPath to get relevant SVGs in this column svgs = div.xpath('.//svg[contains(@class, "svg--34293") and @name and @data-type="roadItem"]') for svg in svgs: name = svg.get("name") first_part = name.split()[0] if first_part in {"Banker", "Player"}: road_result_container[current_column].append(first_part) return road_result_container
Key Takeaways
- Quick Win: Switch to the
lxmlparser with BeautifulSoup—this alone can reduce parsing time significantly. - Reduce DOM Hits: Use combined selectors to avoid unnecessary traversals.
- Skip Abstraction: For maximum speed, use
lxmldirectly instead of BeautifulSoup, as it cuts out the middle layer.
Content of the question originates from Stack Exchange, asked by Kate shim

