使用requests与BeautifulSoup爬取波斯语页面后,如何将杂乱文本转为字典?
解决波斯语页面爬取的空格问题并转换为字典格式
一、清理波斯语文本中的冗余空格
波斯语属于RTL(从右到左)语言,页面里的随机空格常包含全角空格、非断行空格( )或连续空白字符,普通strip()无法彻底处理。用正则匹配所有Unicode空白字符统一清理:
import re def clean_persian_text(text): # 替换所有连续空白字符(含全角、半角、换行)为单个空格 cleaned = re.sub(r'\s+', ' ', text) # 去除首尾空格+处理HTML转义的非断行空格 cleaned = cleaned.strip().replace(' ', ' ') return cleaned
二、提取结构化数据并转为字典
根据rightContent下子div的实际结构,分两种常见情况处理:
完整代码示例
import requests from bs4 import BeautifulSoup import re def clean_persian_text(text): cleaned = re.sub(r'\s+', ' ', text) cleaned = cleaned.strip().replace(' ', ' ') return cleaned # 替换为你的目标页面URL url = "https://example.com/target-page" response = requests.get(url) # 强制指定UTF-8编码适配波斯语,避免乱码 response.encoding = 'utf-8' soup = BeautifulSoup(response.text, 'html.parser') right_content = soup.find('div', class_='rightContent') child_divs = right_content.find_all('div') # 情况1:子div按「键-值」交替排列 result_dict = {} if len(child_divs) % 2 == 0: for i in range(0, len(child_divs), 2): key = clean_persian_text(child_divs[i].get_text()) value = clean_persian_text(child_divs[i+1].get_text()) result_dict[key] = value # 情况2:每个子div内部含明确键值标签(如带类名的span) # result_dict = {} # for div in child_divs: # key_elem = div.find('span', class_='persian-key') # value_elem = div.find('span', class_='persian-value') # if key_elem and value_elem: # key = clean_persian_text(key_elem.get_text()) # value = clean_persian_text(value_elem.get_text()) # result_dict[key] = value print(result_dict)
额外注意事项
- 若波斯语仍乱码,可替换
response.encoding = 'utf-8'为response.encoding = response.apparent_encoding自动检测编码。 - 若遇到特殊波斯语字符被误删,可调整正则保留指定字符:
re.sub(r'[^\w\u0600-\u06FF\s]+', ' ', text)(仅保留波斯语字符、字母数字和空格)。
内容的提问来源于stack exchange,提问作者user20884512
相关产品推荐
相关产品推荐

