循环爬取黄页数据时,Beautiful Soup提取的文本无法追加至Python DataFrame的TypeError问题
解决向DataFrame追加字符串触发的TypeError问题
嘿,我来帮你搞定这个问题!你遇到的TypeError: cannot concatenate object of type '<class 'str'>'; only Series and DataFrame objs are valid其实很好解释——Pandas里DataFrame的列本质是Series对象,它的append()方法只接受其他Series或DataFrame,你直接传字符串肯定会报错。
下面给你两种靠谱的解决方案,优先推荐第一种(效率更高):
方案一:先收集数据到普通列表,最后一次性生成DataFrame
这种方法是Pandas处理批量数据的常规操作,尤其是数据量较大时,比循环追加DataFrame高效太多。步骤很简单:
- 先初始化两个空列表,分别用来存企业名称和提取到的电话号码
- 循环遍历URL时,把有效数据append到列表里
- 循环结束后,用两个列表直接生成DataFrame
修改后的完整代码如下:
import requests from bs4 import BeautifulSoup import pandas as pd # 初始化空列表存储数据 business_names = [] phones = [] # 假设sc_df是你的源DataFrame,这里用循环遍历每一行 for i in range(len(sc_df)): try: # 构造请求URL url = f"https://www.yellowpages.com/search?search_terms={sc_df['Business Name'][i]}&geo_location_terms={sc_df['City'][i]}+{sc_df['State'][i]}" response = requests.get(url) soup = BeautifulSoup(response.text, 'html.parser') # 提取电话号码 phone_element = soup.find('div', class_='info').find('div', class_='phones phone primary') if phone_element is not None: phone_text = phone_element.text.strip() # 去除多余空格 print(phone_text) # 把数据加入列表 business_names.append(sc_df['Business Name'][i]) phones.append(phone_text) else: print("未找到电话号码") # 可选:记录未找到的情况,用None填充 business_names.append(sc_df['Business Name'][i]) phones.append(None) except requests.exceptions.RequestException as e: print(f"请求出错:{e}") # 出错时也记录对应企业和空值 business_names.append(sc_df['Business Name'][i]) phones.append(None) # 最后一次性生成目标DataFrame new_number_df = pd.DataFrame({ 'Business Name': business_names, 'Phone': phones })
方案二:循环中追加单行DataFrame(不推荐,效率低)
如果你一定要在循环里实时追加到DataFrame,可以每次创建一个单行DataFrame,然后用pd.concat()合并到原DataFrame中。注意要加ignore_index=True避免索引混乱:
import requests from bs4 import BeautifulSoup import pandas as pd # 初始化空的目标DataFrame new_number_df = pd.DataFrame(columns=['Business Name', 'Phone']) for i in range(len(sc_df)): try: response = requests.get(f"https://www.yellowpages.com/search?search_terms={sc_df['Business Name'][i]}&geo_location_terms={sc_df['City'][i]}+{sc_df['State'][i]}") soup = BeautifulSoup(response.text, 'html.parser') test = soup.find('div', attrs={'class':'info'}).find('div', attrs={'class':'phones phone primary'}) if test is not None: text = test.text.strip() print(text) # 创建单行DataFrame new_row = pd.DataFrame({ 'Business Name': [sc_df['Business Name'][i]], 'Phone': [text] }) # 合并到原DataFrame new_number_df = pd.concat([new_number_df, new_row], ignore_index=True) else: print(None) except requests.exceptions.RequestException as e: raise SystemExit(e)
为什么优先选方案一?
因为Pandas的DataFrame是基于Numpy数组实现的,每次用concat追加都会重新创建整个DataFrame对象,数据量越大,这种方式越慢。而先存列表再一次性生成DataFrame,是直接在内存中构建数据结构,效率会高很多。
内容的提问来源于stack exchange,提问作者ryan muir
相关产品推荐
相关产品推荐

