按亮度类别在球面上均匀分布恒星的技术咨询
恒星球面映射的亮度类别均匀分布实现问题
我正在开展一个将恒星映射到虚拟球面的项目,要求恒星彼此之间及与球心(0,0,0)均匀间隔,且球面各区域需按指定比例分布不同亮度类别的恒星,确保整个球面覆盖均匀。
现有思路与问题
- 归一化处理:已将恒星位置向量缩放到固定半径,完成球面归一化。
- 均匀点生成:使用Fibonacci球体采样法生成了1000个球面上均匀分布的点。
- 初始匹配:不考虑亮度时,将每个点匹配到最近恒星的方案效果良好。
- 类别分布难题:按指定亮度类别比例分配时,难以同时保证类别比例与恒星均匀分布。
- 当前尝试的局限:随机打乱生成点后按类别分配的方法,无法实现理想的均匀分布,尤其是类别'8'的恒星分布仍不理想。
可行解决方案
1. 分层球面分区分配法
将球面划分为多个大小均等的区域(比如八面体细分、经纬度网格或基于Fibonacci采样的分组),在每个区域内按照指定比例分配各亮度类别的恒星配额。从空间分区层面保证每个区域都符合比例,避免局部聚类。
具体步骤:
- 对生成的1000个Fibonacci点进行空间聚类(如K-Means,聚类数设为20-50个分区),确保每个分区的点数量大致均等。
- 每个分区内按照
desired_distribution计算各类别需要分配的恒星数量。 - 在每个分区内,为对应的点匹配该类别下最近的恒星。
2. 带空间约束的贪心匹配法
优先为稀有类别(比如占比19%的'8'类)分配位置,保证其均匀分布,再依次处理其他类别,填充剩余位置时避开已分配恒星的邻近区域。
具体步骤:
- 优先处理占比最低的类别:从Fibonacci点集中为'8'类恒星挑选间隔最大的点集(最远点优先采样),确保这些点均匀分布在球面上。
- 为这些点匹配最近的'8'类恒星。
- 剩余点按比例分配给其他类别,匹配时跳过已分配恒星的邻近区域,避免局部聚集。
3. 基于类别权重的K近邻匹配
修改匹配逻辑,为每个Fibonacci点计算所有类别恒星的最近邻,然后根据全局比例随机选择符合要求的类别,同时加入空间惩罚项——如果某区域已分配过多同类别恒星,降低该类别被选中的概率。
修改后的实现代码(带空间约束的贪心匹配法)
import pandas as pd import numpy as np from sklearn.neighbors import NearestNeighbors from scipy.spatial.distance import cdist def classify_stars(vmag): if vmag >= 10: return '1' elif 7 <= vmag < 10: return '3' elif 6 <= vmag < 7: return '5' elif 3 <= vmag < 6: return '8' elif 1 <= vmag < 3: return '9' else: return '10' def generate_sphere_points(samples=1000, radius=50): points = [] dphi = np.pi * (3. - np.sqrt(5.)) # 黄金角近似值(弧度) for i in range(samples): y = 1 - (i / float(samples - 1)) * 2 # y范围从1到-1 r = np.sqrt(1 - y * y) # y处的截面半径 theta = dphi * i # 黄金角增量 x = np.cos(theta) * r z = np.sin(theta) * r points.append((x * radius, y * radius, z * radius)) return np.array(points) # 数据加载与预处理 df = pd.read_csv('hygdata_v3.csv', usecols=['hip', 'x', 'y', 'z', 'mag']) df.dropna(subset=['hip', 'x', 'y', 'z', 'mag'], inplace=True) df['hip'] = df['hip'].astype(int) df['norm'] = np.sqrt(df['x']**2 + df['y']**2 + df['z']**2) df['x'] = 50 * df['x'] / df['norm'] df['y'] = 50 * df['y'] / df['norm'] df['z'] = 50 * df['z'] / df['norm'] df.drop(columns='norm', inplace=True) df['class'] = df['mag'].apply(classify_stars) points = generate_sphere_points(samples=1000) desired_distribution = {'1': 0.27, '3': 0.27, '5': 0.27, '8': 0.19, '9': 0, '10': 0} total_points = len(points) category_points = {k: int(v * total_points) for k, v in desired_distribution.items()} # 优先处理类别'8',确保均匀分布 sampled_df = pd.DataFrame() remaining_points = points.copy() # 步骤1:为类别'8'挑选均匀分布的点 cat8_count = category_points['8'] if cat8_count > 0: # 最远点优先采样,保证点的均匀分布 selected_indices = [] # 随机选第一个点 current_idx = np.random.choice(len(remaining_points)) selected_indices.append(current_idx) for _ in range(cat8_count - 1): # 计算剩余点到已选点的最小距离 distances = cdist(remaining_points, remaining_points[selected_indices]).min(axis=1) # 选距离最大的点 current_idx = np.argmax(distances) selected_indices.append(current_idx) cat8_points = remaining_points[selected_indices] # 匹配最近的'8'类恒星 cat8_stars = df[df['class'] == '8'] nbrs = NearestNeighbors(n_neighbors=1).fit(cat8_stars[['x', 'y', 'z']]) _, indices = nbrs.kneighbors(cat8_points) assigned_cat8 = cat8_stars.iloc[indices.flatten()] sampled_df = pd.concat([sampled_df, assigned_cat8], ignore_index=True) # 移除已分配的点 remaining_points = np.delete(remaining_points, selected_indices, axis=0) # 步骤2:处理其他类别,按比例分配剩余点 remaining_cats = [k for k in desired_distribution if k != '8' and category_points[k] > 0] for category in remaining_cats: count = category_points[category] # 从剩余点中取对应数量的点(保持原有Fibonacci均匀性) take_count = min(count, len(remaining_points)) cat_points = remaining_points[:take_count] # 匹配最近的该类别恒星 cat_stars = df[df['class'] == category] nbrs = NearestNeighbors(n_neighbors=1).fit(cat_stars[['x', 'y', 'z']]) _, indices = nbrs.kneighbors(cat_points) assigned_stars = cat_stars.iloc[indices.flatten()] sampled_df = pd.concat([sampled_df, assigned_stars], ignore_index=True) # 移除已分配的点 remaining_points = remaining_points[take_count:] # 去重并输出结果 sampled_df = sampled_df.drop_duplicates(subset='hip') category_8_stars = sampled_df[sampled_df['class'] == '8'] hip_values_category_8 = category_8_stars['hip'].astype(str).tolist() print(hip_values_category_8)
内容的提问来源于stack exchange,提问作者Clms
相关产品推荐
相关产品推荐

