如何优化重复的elements_by_id元素格式化代码?
Hey there! Let's tackle that repetitive code and get your element parsing working smoothly. First, let's figure out why your initial parse_element function didn't work—then we'll build a cleaner, more efficient solution.
Why Your Original parse_element Failed
The issue is straightforward: your function modifies the variable x internally but doesn't return the processed value. In Python, reassigning a parameter inside a function doesn't change the original variable outside of it. So when you called parse_element(description), the result was just discarded instead of being assigned back to your variable.
Fixed Basic parse_element
First, let's fix that core function by adding a return statement:
def parse_element(elements): # Extract text from each element and join into a single string return " ".join([elem.text for elem in elements])
Now you can use it correctly in your main function:
description_elements = driver.find_elements_by_id('some-id') description = parse_element(description_elements)
Take It Further: Combine Finding and Parsing
Even better, we can eliminate more repetition by creating a function that handles both finding the elements and parsing their text. This way you don't have to write driver.find_elements_* and then call parse_element every time.
We can even add an optional parameter for cases like your location field, where you need to append extra text:
def get_parsed_element(driver, locator_type, locator_value, suffix=""): # Find elements based on the locator type if locator_type == "id": elements = driver.find_elements_by_id(locator_value) elif locator_type == "xpath": elements = driver.find_elements_by_xpath(locator_value) # Add more locator types (like class_name, css_selector) if needed # Parse and append suffix if provided parsed_text = " ".join([elem.text for elem in elements]) return parsed_text + suffix
Optimized Main Function
Now your get_description function becomes super clean and easy to read:
def get_description(driver, links): for link in links: # Don't forget to navigate to the link first (I assume this step was missing in your snippet) # driver.get(link) description = get_parsed_element(driver, "id", "some-id") title = get_parsed_element(driver, "id", "different-id") company = get_parsed_element(driver, "id", "another-different-id") location = get_parsed_element(driver, "id", "location-id", suffix=" United Kingdom") salary = get_parsed_element(driver, "xpath", "//*[@id='randomly generated id']/div[3]/span[1]") # Do something with these variables (like append to a list, save to a database, etc.)
Bonus: Batch Processing for Even More Flexibility
If you ever need to add more fields later, you can define a list of field configurations and loop through them—this makes scaling and updating your code even easier:
def get_description(driver, links): field_configs = [ ("description", "id", "some-id", ""), ("title", "id", "different-id", ""), ("company", "id", "another-different-id", ""), ("location", "id", "location-id", " United Kingdom"), ("salary", "xpath", "//*[@id='randomly generated id']/div[3]/span[1]", ""), ] for link in links: # Navigate to the target link # driver.get(link) # Extract all fields in one loop extracted_data = {} for field_name, loc_type, loc_val, suffix in field_configs: extracted_data[field_name] = get_parsed_element(driver, loc_type, loc_val, suffix) # Use the extracted data (e.g., print(extracted_data), save to a CSV file)
This approach keeps your code DRY (Don't Repeat Yourself), easier to maintain, and simpler to update if you need to add or modify fields later.
内容的提问来源于stack exchange,提问作者lawson

