You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

按亮度类别在球面上均匀分布恒星的技术咨询

恒星球面映射的亮度类别均匀分布实现问题

我正在开展一个将恒星映射到虚拟球面的项目,要求恒星彼此之间及与球心(0,0,0)均匀间隔,且球面各区域需按指定比例分布不同亮度类别的恒星,确保整个球面覆盖均匀。

现有思路与问题

  • 归一化处理:已将恒星位置向量缩放到固定半径,完成球面归一化。
  • 均匀点生成:使用Fibonacci球体采样法生成了1000个球面上均匀分布的点。
  • 初始匹配:不考虑亮度时,将每个点匹配到最近恒星的方案效果良好。
  • 类别分布难题:按指定亮度类别比例分配时,难以同时保证类别比例与恒星均匀分布。
  • 当前尝试的局限:随机打乱生成点后按类别分配的方法,无法实现理想的均匀分布,尤其是类别'8'的恒星分布仍不理想。

可行解决方案

1. 分层球面分区分配法

将球面划分为多个大小均等的区域(比如八面体细分、经纬度网格或基于Fibonacci采样的分组),在每个区域内按照指定比例分配各亮度类别的恒星配额。从空间分区层面保证每个区域都符合比例,避免局部聚类。

具体步骤:

  • 对生成的1000个Fibonacci点进行空间聚类(如K-Means,聚类数设为20-50个分区),确保每个分区的点数量大致均等。
  • 每个分区内按照desired_distribution计算各类别需要分配的恒星数量。
  • 在每个分区内,为对应的点匹配该类别下最近的恒星。

2. 带空间约束的贪心匹配法

优先为稀有类别(比如占比19%的'8'类)分配位置,保证其均匀分布,再依次处理其他类别,填充剩余位置时避开已分配恒星的邻近区域。

具体步骤:

  1. 优先处理占比最低的类别:从Fibonacci点集中为'8'类恒星挑选间隔最大的点集(最远点优先采样),确保这些点均匀分布在球面上。
  2. 为这些点匹配最近的'8'类恒星。
  3. 剩余点按比例分配给其他类别,匹配时跳过已分配恒星的邻近区域,避免局部聚集。

3. 基于类别权重的K近邻匹配

修改匹配逻辑,为每个Fibonacci点计算所有类别恒星的最近邻,然后根据全局比例随机选择符合要求的类别,同时加入空间惩罚项——如果某区域已分配过多同类别恒星,降低该类别被选中的概率。

修改后的实现代码(带空间约束的贪心匹配法)

import pandas as pd
import numpy as np
from sklearn.neighbors import NearestNeighbors
from scipy.spatial.distance import cdist

def classify_stars(vmag):
    if vmag >= 10:
        return '1'
    elif 7 <= vmag < 10:
        return '3'
    elif 6 <= vmag < 7:
        return '5'
    elif 3 <= vmag < 6:
        return '8'
    elif 1 <= vmag < 3:
        return '9'
    else:
        return '10'

def generate_sphere_points(samples=1000, radius=50):
    points = []
    dphi = np.pi * (3. - np.sqrt(5.))  # 黄金角近似值(弧度)
    for i in range(samples):
        y = 1 - (i / float(samples - 1)) * 2  # y范围从1到-1
        r = np.sqrt(1 - y * y)  # y处的截面半径

        theta = dphi * i  # 黄金角增量
        x = np.cos(theta) * r
        z = np.sin(theta) * r

        points.append((x * radius, y * radius, z * radius))
    return np.array(points)

# 数据加载与预处理
df = pd.read_csv('hygdata_v3.csv', usecols=['hip', 'x', 'y', 'z', 'mag'])
df.dropna(subset=['hip', 'x', 'y', 'z', 'mag'], inplace=True)
df['hip'] = df['hip'].astype(int)
df['norm'] = np.sqrt(df['x']**2 + df['y']**2 + df['z']**2)
df['x'] = 50 * df['x'] / df['norm']
df['y'] = 50 * df['y'] / df['norm']
df['z'] = 50 * df['z'] / df['norm']
df.drop(columns='norm', inplace=True)
df['class'] = df['mag'].apply(classify_stars)

points = generate_sphere_points(samples=1000)
desired_distribution = {'1': 0.27, '3': 0.27, '5': 0.27, '8': 0.19, '9': 0, '10': 0}
total_points = len(points)
category_points = {k: int(v * total_points) for k, v in desired_distribution.items()}

# 优先处理类别'8',确保均匀分布
sampled_df = pd.DataFrame()
remaining_points = points.copy()

# 步骤1:为类别'8'挑选均匀分布的点
cat8_count = category_points['8']
if cat8_count > 0:
    # 最远点优先采样,保证点的均匀分布
    selected_indices = []
    # 随机选第一个点
    current_idx = np.random.choice(len(remaining_points))
    selected_indices.append(current_idx)
    
    for _ in range(cat8_count - 1):
        # 计算剩余点到已选点的最小距离
        distances = cdist(remaining_points, remaining_points[selected_indices]).min(axis=1)
        # 选距离最大的点
        current_idx = np.argmax(distances)
        selected_indices.append(current_idx)
    
    cat8_points = remaining_points[selected_indices]
    # 匹配最近的'8'类恒星
    cat8_stars = df[df['class'] == '8']
    nbrs = NearestNeighbors(n_neighbors=1).fit(cat8_stars[['x', 'y', 'z']])
    _, indices = nbrs.kneighbors(cat8_points)
    assigned_cat8 = cat8_stars.iloc[indices.flatten()]
    sampled_df = pd.concat([sampled_df, assigned_cat8], ignore_index=True)
    
    # 移除已分配的点
    remaining_points = np.delete(remaining_points, selected_indices, axis=0)

# 步骤2:处理其他类别,按比例分配剩余点
remaining_cats = [k for k in desired_distribution if k != '8' and category_points[k] > 0]
for category in remaining_cats:
    count = category_points[category]
    # 从剩余点中取对应数量的点(保持原有Fibonacci均匀性)
    take_count = min(count, len(remaining_points))
    cat_points = remaining_points[:take_count]
    # 匹配最近的该类别恒星
    cat_stars = df[df['class'] == category]
    nbrs = NearestNeighbors(n_neighbors=1).fit(cat_stars[['x', 'y', 'z']])
    _, indices = nbrs.kneighbors(cat_points)
    assigned_stars = cat_stars.iloc[indices.flatten()]
    sampled_df = pd.concat([sampled_df, assigned_stars], ignore_index=True)
    # 移除已分配的点
    remaining_points = remaining_points[take_count:]

# 去重并输出结果
sampled_df = sampled_df.drop_duplicates(subset='hip')
category_8_stars = sampled_df[sampled_df['class'] == '8']
hip_values_category_8 = category_8_stars['hip'].astype(str).tolist()
print(hip_values_category_8)

内容的提问来源于stack exchange,提问作者Clms

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 22:42:04