如何从网页HTML提取文本作为变量,应用到其他输入场景?
网页信息提取逻辑实现
修改后的完整代码
from selenium.common.exceptions import NoSuchElementException from selenium.webdriver.common.by import By import time from selenium.webdriver.chrome.options import Options from selenium.webdriver.chrome.service import Service from webdriver_manager.chrome import ChromeDriverManager from selenium import webdriver import re def ClientInformation(): CI = textbox.get('1.0', END) # 按空格拆分客户信息 CIS = re.split(r' ', str(CI)) FN = CIS[0] LN = CIS[1] CT = CIS[2] ST = CIS[3] options = Options() options.add_experimental_option("detach", True) # 启动Chrome并打开目标搜索页面 driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()), options=options) driver.get(f"https://www.truepeoplesearch.com/results?name={FN}%20{LN}&citystatezip={CT}%20{ST}") time.sleep(1) # 延长等待时间,0.05秒太短,页面大概率未加载完成 # 提取电话号码,不存在则赋值为N/A try: phone_number = driver.find_element(By.CLASS_NAME, "phone-link").text # 替换为实际网页中电话号码元素的选择器 except NoSuchElementException: phone_number = "N/A" # 提取"smaller"内容,不存在则赋值为N/A try: smaller_content = driver.find_element(By.CLASS_NAME, "smaller").text # 替换为实际网页中对应元素的选择器 except NoSuchElementException: smaller_content = "N/A" # 这里可将变量导入到指定位置,比如写入文本框、传入其他函数等 print(f"电话号码: {phone_number}") print(f"smaller内容: {smaller_content}") # 如需返回变量供外部使用,可添加return语句 # return phone_number, smaller_content
关键注意事项
- 等待时间:必须把
time.sleep(0.05)改成1秒以上,网页加载需要时间,过短的等待会导致元素未渲染就触发查找,大概率报错。 - 元素选择器:代码中的
phone-link和smaller是假设的类名,你需要根据实际HTML代码替换成正确的选择器。比如元素用id标识就用By.ID, "元素id",用CSS选择器就用By.CSS_SELECTOR, "对应CSS路径"。 - 异常处理:通过
try-except捕获NoSuchElementException,确保元素不存在时程序不崩溃,而是自动赋值为N/A。 - 变量使用:提取后的
phone_number和smaller_content就是你需要的变量,直接在后续代码中调用即可,比如写入目标位置或存储。
内容的提问来源于stack exchange,提问作者IIzzy
相关产品推荐
相关产品推荐

