如何通过Python脚本高效安全获取Active Directory用户的递归组
Let's tackle why your current code is running too slow, walk through your existing setup, and share actionable optimizations to speed things up.
Your Current Logic
Your workflow makes sense for ensuring you get the most up-to-date group data, but the implementation is causing unnecessary overhead:
- Fetch a user's details by querying all domain controllers, then pick the latest entry via
whenChangedto get direct groups (memberOf) - Recursively repeat this DC-wide query for each group to fetch its parent subgroups
- Track queried groups to avoid duplicates and infinite loops (e.g., Group A ↔ Group B mutual memberships)
Your Configuration
{ "domain_controllers": ["abc.domainname.companyname.com", "pqr.domainname.companyname.com"], "name": "domainname", "user": "*", "password": "*", "base_dn": "DC=domainname,DC=companyname,DC=com" }
Your Current (Slow) Code
Note: Fixed a typo (serahc_base → search_base) and kept placeholder markers for your actual connection/conversion logic:
def get_latest_entry(self, domain_name, data_type, ldap_query): domain = self.config.get_domain(domain_name) ldap_attrs = ["memberOf", "whenChanged"] ldap_entries_for_all_dcs = [] for hostname in domain.domain_controllers: with "<<ldap_connection>>": # Replace with your actual LDAP connection setup try: ldap_connection.search(search_base=domain.base_dn, search_filter=ldap_query, attributes=ldap_attrs) except ldap3.core.exceptions.LDAPInvalidFilterError: return None if not ldap_connection.entries: continue json_entry = <<self.converttojson>>(ldap_connection.entries[0]) # Replace with your conversion method if json_entry["memberOf"]: json_entry["memberOf"] = [str(memberOf.encode('utf-8').decode('latin-1')) for memberOf in json_entry["memberOf"]] ldap_entries_for_all_dcs.append(json_entry) return max(ldap_entries_for_all_dcs, key=lambda ldap_entry: ldap_entry['whenChanged'])["memberOf"] if ldap_entries_for_all_dcs else None def get_user_groups(self, sam_account_name, domain_name): return self.get_latest_entry( domain_name=domain_name, data_type="user", ldap_query="(&(objectCategory=Person)(objectClass=user)(sAMAccountName={}))".format(sam_account_name) ) def get_sub_groups(self, distinguished_name, domain_name): return self.get_latest_entry( domain_name=domain_name, data_type="group", ldap_query="(&(objectCategory=Group)(objectClass=Group)(distinguishedName={}))".format(distinguished_name) ) def get_recursive_groups(self, sam_account_name, domain_name): def nestedgrouplookup(distinguished_name, domain_name): if distinguished_name not in recursive_groups_list: recursive_groups_list.append(distinguished_name) nested_list = self.get_sub_groups(distinguished_name, domain_name) if nested_list: for memberOf in nested_list: nestedgrouplookup(memberOf, domain_name) recursive_groups_list = [] user_groups_list = self.get_user_groups(sam_account_name, domain_name) if not user_groups_list: return [] for distinguished_name in user_groups_list: nestedgrouplookup(distinguished_name, domain_name) return recursive_groups_list
Key Performance Bottlenecks & Fixes
1. Querying All DCs for Every Single Entry
Problem: Every call to get_latest_entry connects to all DCs just to compare whenChanged timestamps. Connection setup/teardown and redundant queries add massive overhead, especially during recursion.
Fix:
- Pick a single preferred DC (e.g., a read-only replica or the closest one) and only fall back to others if it fails. AD replication ensures most DCs have up-to-date data within minutes.
- Reuse LDAP connections instead of creating a new one per query. Connection pooling is even better for high-volume use cases.
2. Recursive Per-Group Queries
Problem: For deep group hierarchies, you're making dozens (or hundreds) of individual LDAP calls. This is inefficient.
Fix:
Use Active Directory's built-in LDAP_MATCHING_RULE_IN_CHAIN filter to fetch all recursive groups in one query. This filter traverses group membership hierarchies automatically.
3. Unnecessary Encoding/Decoding
Problem: Manually encoding/decoding memberOf values adds extra processing time.
Fix:
Configure your LDAP connection to use UTF-8 encoding by default, so you don't need to handle this step for every entry.
Optimized Code Example
This version cuts runtime drastically by leveraging AD's native recursive lookup and reducing redundant DC queries:
def get_recursive_groups_optimized(self, sam_account_name, domain_name): domain = self.config.get_domain(domain_name) # Use a single DC (add failover logic if needed) primary_dc = domain.domain_controllers[0] recursive_groups = set() # Automatically avoids duplicates with "<<ldap_connection>>": # Reuse this connection for both queries # Step 1: Get the user's distinguished name user_filter = "(&(objectCategory=Person)(objectClass=user)(sAMAccountName={}))".format(sam_account_name) ldap_connection.search( search_base=domain.base_dn, search_filter=user_filter, attributes=["distinguishedName"] ) if not ldap_connection.entries: return [] user_dn = str(ldap_connection.entries[0].distinguishedName) # Step 2: Use LDAP_MATCHING_RULE_IN_CHAIN to get all recursive groups recursive_group_filter = "(member:1.2.840.113556.1.4.1941:={})".format(user_dn) ldap_connection.search( search_base=domain.base_dn, search_filter="(&(objectCategory=Group)(objectClass=Group){})".format(recursive_group_filter), attributes=["distinguishedName"] ) # Collect all group DNs for entry in ldap_connection.entries: recursive_groups.add(str(entry.distinguishedName)) return list(recursive_groups)
Extra Optimizations
- Cache Results: Cache group memberships for frequently accessed users/groups to avoid re-querying.
- Batch Queries: If processing multiple users, batch their lookups instead of querying one-by-one.
- Consistency Check: If you must verify
whenChanged, query a single DC and check if the entry is newer than your last cached timestamp instead of querying all DCs.
内容的提问来源于stack exchange,提问作者SmiP

