使用Threading爬取YellowPages时,locations列表仅保留最后一个州数据
问题:多线程爬取Yellow Pages仅获取最后一个州的数据
我尝试从yellowpages.com爬取数据,该网站通过URL https://www.yellowpages.com/state-<state-abbreviation>?page=<letter> 展示某州内以指定字母开头的城市列表,例如纽约州以字母'c'开头的城市对应URL为https://www.yellowpages.com/state-ny?page=c。
需求是将所有“城市,州”组合存入locations变量并写入文件,因单线程效率低改用多线程实现。运行后日志显示完成全部1300个页面(50个州×26个字母)的请求,但文件中仅写入最后一个州怀俄明州(Wyoming)的A-Z城市数据,其他州数据缺失。
我的代码:
def get_session(): if not hasattr(thread_local, 'session'): thread_local.session = requests.Session() return thread_local.session def download_site(url): """ Make request to url and scrape data using bs4""" session = get_session() with session.get(url) as response: logging.info(f"Read {len(response.content)} from {url}") scrape_data(response) def download_all_sites(urls): """ call download_site() on list of urls""" with concurrent.futures.ThreadPoolExecutor(max_workers = 50) as executor: executor.map(download_site, urls) def scrape_data(response): """uses bs4 to get city, state combo from yellowpages html and appends to global locations list""" soup = BeautifulSoup(response.text, 'html.parser') ul_elements = soup.find_all('ul') for ul_element in ul_elements: anchor_elements = ul_element.find_all('a') for element in anchor_elements: locations.append(element.text + ',' + state_abbrieviated) if __name__ == '__main__': logging.basicConfig(level=logging.INFO) urls = [] # will hold yellowpages urls locations = [] # will hold scraped 'city, state' combinations, modified by scrape_data() function states = { 'AK': 'Alaska', 'AL': 'Alabama', 'AR': 'Arkansas', 'AZ': 'Arizona', 'CA': 'California', 'CO': 'Colorado', 'CT': 'Connecticut', 'DC': 'District of Columbia', 'DE': 'Delaware', 'FL': 'Florida', 'GA': 'Georgia', 'HI': 'Hawaii', 'IA': 'Iowa', 'ID': 'Idaho', 'IL': 'Illinois', 'IN': 'Indiana', 'KS': 'Kansas', 'KY': 'Kentucky', 'LA': 'Louisiana', 'MA': 'Massachusetts', 'MD': 'Maryland', 'ME': 'Maine', 'MI': 'Michigan', 'MN': 'Minnesota', 'MO': 'Missouri', 'MS': 'Mississippi', 'MT': 'Montana', 'NC': 'North Carolina', 'ND': 'North Dakota', 'NE': 'Nebraska', 'NH': 'New Hampshire', 'NJ': 'New Jersey', 'NM': 'New Mexico', 'NV': 'Nevada', 'NY': 'New York', 'OH': 'Ohio', 'OK': 'Oklahoma', 'OR': 'Oregon', 'PA': 'Pennsylvania', 'RI': 'Rhode Island', 'SC': 'South Carolina', 'SD': 'South Dakota', 'TN': 'Tennessee', 'TX': 'Texas', 'UT': 'Utah', 'VA': 'Virginia', 'VT': 'Vermont', 'WA': 'Washington', 'WI': 'Wisconsin', 'WV': 'West Virginia', 'WY': 'Wyoming' } letters = ['a','b','c','d','e','f','g','h','i','j','k','l','m','n','o', 'p','q','r','s','t','u','v','w','x','y','z'] # build list of urls that need to be scrape for state_abbrieviated ,state_full in states.items(): for letter in letters: url = f'https://www.yellowpages.com/state-{state_abbrieviated}?page={letter}' urls.append(url) # scrape data download_all_sites(urls) logging.info(f"\tSent/Retrieved {len(urls)} requests/responses in {duration} seconds") # write data to file with open('locations.txt','w') as file: for location in locations: file.write(location + '\n')
问题根源
- 全局变量竞态条件:
state_abbrieviated是全局变量,构建URL的循环会快速遍历所有州,最终这个变量会被固定为最后一个州的缩写(WY)。多线程并行执行时,所有线程调用scrape_data都会使用这个最终值,导致所有城市错误关联到怀俄明州。 - 代码缩进错误:原代码中
download_all_sites调用、日志输出和文件写入部分存在缩进错误,会导致运行异常(不过你提到日志显示完成所有请求,可能是粘贴时的格式问题)。
修复方案
方案1:绑定URL与州缩写,传递给下载函数
让每个请求携带对应的州缩写,避免依赖全局变量:
def download_site(url, state_abbr): """ Make request to url and scrape data using bs4""" session = get_session() with session.get(url) as response: logging.info(f"Read {len(response.content)} from {url}") scrape_data(response, state_abbr) def download_all_sites(url_state_pairs): """ call download_site() on list of (url, state_abbr) pairs""" with concurrent.futures.ThreadPoolExecutor(max_workers=50) as executor: executor.map(lambda x: download_site(*x), url_state_pairs) def scrape_data(response, state_abbr): """uses bs4 to get city, state combo from yellowpages html and appends to global locations list""" soup = BeautifulSoup(response.text, 'html.parser') ul_elements = soup.find_all('ul') for ul_element in ul_elements: anchor_elements = ul_element.find_all('a') for element in anchor_elements: locations.append(f"{element.text},{state_abbr}") # 构建URL和州缩写的配对列表 url_state_pairs = [] for state_abbrieviated ,state_full in states.items(): for letter in letters: url = f'https://www.yellowpages.com/state-{state_abbrieviated}?page={letter}' url_state_pairs.append((url, state_abbrieviated)) # 调用修改后的下载函数 download_all_sites(url_state_pairs)
方案2:从URL中解析州缩写
如果不想修改函数参数,可以直接从响应的URL中提取州缩写:
def scrape_data(response): """uses bs4 to get city, state combo from yellowpages html and appends to global locations list""" # 从URL中提取州缩写 state_abbr = response.url.split('state-')[1].split('?')[0] soup = BeautifulSoup(response.text, 'html.parser') ul_elements = soup.find_all('ul') for ul_element in ul_elements: anchor_elements = ul_element.find_all('a') for element in anchor_elements: locations.append(f"{element.text},{state_abbr}")
额外修复:修正代码缩进
确保download_all_sites调用、日志输出和文件写入部分的缩进与if __name__ == '__main__':下的代码对齐:
# 修正后的主函数部分 if __name__ == '__main__': logging.basicConfig(level=logging.INFO) urls = [] locations = [] # ... 省略states和letters定义 ... # build list of urls that need to be scrape url_state_pairs = [] for state_abbrieviated ,state_full in states.items(): for letter in letters: url = f'https://www.yellowpages.com/state-{state_abbrieviated}?page={letter}' url_state_pairs.append((url, state_abbrieviated)) # scrape data import time start_time = time.time() download_all_sites(url_state_pairs) duration = time.time() - start_time logging.info(f"\tSent/Retrieved {len(url_state_pairs)} requests/responses in {duration:.2f} seconds") # write data to file with open('locations.txt','w') as file: for location in locations: file.write(location + '\n')
内容的提问来源于stack exchange,提问作者Justin
相关产品推荐
相关产品推荐

