3.5 KiB
3.5 KiB
title, created, updated, type, tags, arxiv, authors, venue, year, sources
| title | created | updated | type | tags | arxiv | authors | venue | year | sources | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| CuBAS: Information Geometric Curvature-Based Adaptive Sampling for Supervised Classification | 2026-07-10 | 2026-07-10 | paper |
|
2607.03145 |
|
arXiv preprint | 2026 |
|
CuBAS: Curvature-Based Adaptive Sampling (2026)
CuBAS(Curvature-Based Adaptive Sampling)是一个信息几何驱动的自适应数据选择框架,将监督分类中的训练集筛选转化为statistical-manifold的局部曲率分析问题。
核心问题
如何在带标签数据集中识别最具信息量的样本?
传统方法依赖随机采样或基于分类器不确定性的试探,CuBAS 则从数据分布的内在几何出发——将带标签 k-NN 图视为由 potts-markov-random-field 诱导的统计流形,曲率直接编码信息密度。
方法论架构
Labeled Dataset D
│
▼
k-NN Graph G ──→ Potts MRF (β by MPL)
│
▼
Node-wise Curvature Scores S_i(β) = -Ψ_i/(Φ_i+λ)
│ ↑
│ Φ_i: 1st-order Fisher (metric)
│ Ψ_i: 2nd-order Fisher (curvature)
▼
Adaptive Threshold ──→ L (low-curvature) + H (high-curvature)
│
▼
Curvature-Aware Subsampling → Compact & Informative Training Set
关键贡献
- 统计流形曲率作为信息量度量:通过 observed-fisher-information 比值定义每个节点的标量曲率
- CuBAS 算法:模型无关、仅需图拓扑 + Potts 充分统计量,无需预训练分类器
- 自适应阈值级联:adaptive-threshold-estimation,每数据集自动调优
- 大规模验证:60+ 数据集(表格/图像/基因组),15 种训练比例 × 100 次重复,全胜 Random 和 Entropy
关键结果
| 场景 | CuBAS | Entropy | Random |
|---|---|---|---|
| 全部 60+ 数据集均值 | 0.8455 | 0.7623 | 0.7273 |
| 10% 训练预算(42 数据集) | 0.8691 | 0.7345 | — |
| breast_cancer | 0.9996 | 0.9752 | 0.9469 |
| wine-quality-red | 0.5214 | 0.2731 | 0.2703 |
Wilcoxon 检验 p < 10⁻¹⁰,证实 CuBAS 优势系统性而非偶然。
数据效率特性
CuBAS 在小训练比例下即可达到 Random 采样在大训练比例下的精度——这一数据效率特性在医学影像、基因组分析等标注成本高的领域有直接应用价值。
限制
- 超小样本高维场景(n ≪ p):k-NN 图稀疏,曲率信号噪声大
- 文本稀疏 bag-of-words 特征:k-NN 图几何结构弱
- 当前仅被动采样,未扩展至主动学习